OpenAI has launched a new mode called Ultrafast intended to speed up its most capable model, GPT-5.6 Sol. In a blog post on Thursday, the company said the feature is designed to substantially reduce turnaround time and increase the amount of useful work the model can do per second.
How fast is it?
OpenAI claims Ultrafast can operate at up to 14× the speed of standard processing and produce as many as 750 output tokens per second. (An "output token" refers to the distinct pieces of text a large language model generates when interacting with a human.)
Suggested use cases
The company recommends Ultrafast for latency-sensitive enterprise workflows, notably:
- incident response,
- customer service and support,
- financial market analysis,
- e-commerce, and related areas.
OpenAI noted that until now achieving real-time speed typically required choosing a smaller or more specialized model; Ultrafast is presented as a step toward delivering more useful work per second with a high-capability model.
Technical backing and availability
The Ultrafast preview is powered by OpenAI's partnership with chipmaker Cerebras. The preview release is currently available only to a small group of customers; OpenAI said it will expand access as capacity grows.
Competitors and context
Rivals have rolled out their own accelerated modes: for example, Anthropic's Claude offers a fast mode, although OpenAI says it does not reach the same speeds Ultrafast aims to deliver.
Why this matters
Faster response rates improve the effectiveness of applications where low latency is critical. By focusing on throughput and speed with a large, general-purpose model, OpenAI's Ultrafast move signals a push to make high-capability models more practical for real-time, enterprise-grade use cases.



