The news
OpenAI announced the preview of Ultrafast mode on August 13. The new tier targets enterprise customers who need higher throughput from the GPT-5.6 Sol model. It delivers up to 14× the speed of the standard version and peaks at 750 output tokens per second.
The change is powered by Cerebras systems. OpenAI positions the offering as a direct response to demand from companies that run large workloads through the API. Access requires an approved enterprise account and routes requests through the existing OpenAI API endpoints rather than any consumer-facing product.
Context
Until now, GPT-5.6 Sol operated at the speeds available on OpenAI’s primary infrastructure. Enterprise teams that required faster responses had to accept lower latency only on smaller models or wait for batch processing. The Ultrafast tier removes that trade-off for the frontier model itself.
The preview is restricted. Access is granted through the OpenAI API rather than the consumer chat interface. Three separate reports confirm the same core specifications: 14× speed, Cerebras hardware, and a 750-token-per-second ceiling. No public pricing or broader rollout timeline appears in the initial announcements.
Details
OpenAI’s own announcement states that Ultrafast mode runs GPT-5.6 Sol at up to 14× the speed of the regular service. The maximum output rate is given as 750 tokens per second. TechCrunch and Neowin both describe the release as a limited preview aimed at enterprise users.
No changes to model weights or training data are mentioned. The performance gain comes from the underlying Cerebras compute layer. The three sources agree on the token-rate figure and the hardware partner; none provide additional benchmarks or pricing details.
The announcement language from OpenAI emphasizes that the tier remains in preview and that availability depends on capacity on the Cerebras cluster. No information is given about how many customers have been granted access or what criteria OpenAI used to select them.
Why it matters
For developers and companies already calling GPT-5.6 Sol, the new tier changes the cost-speed equation. Workloads that previously required multiple smaller models or accepted multi-second delays can now stay on the largest model and finish faster. That shift matters for real-time agents, high-volume summarization, and any pipeline where latency directly affects user experience or operational cost.
The move also signals OpenAI’s willingness to route frontier-model traffic to specialized hardware when standard clusters cannot meet demand. Teams that have built around the model now have a concrete option to increase throughput without switching providers. The preview status means early adopters will test whether the 14× claim holds under sustained load and whether the service remains available once broader access opens.
Because the tier is available only through the API, organizations that rely on the web chat or third-party wrappers will see no immediate change. Those already integrating the model directly can request access and measure the difference against their current latency and cost baselines. The absence of public pricing leaves open the question of whether the speed increase justifies any premium, a point that will become clearer only after more customers gain entry.
---
Sources:
{"word_count": 612, "sources_used": 3, "expanded_sections": ["context", "why_it_matters"]}
No comments yet