Model launches

AI-generated text

OpenAI launches Ultrafast tier—GPT‑5.6 Sol up to 14× faster with Cerebras

OpenAI announced Ultrafast, a new API service tier running GPT‑5.6 Sol that offers up to 14× speed improvement over Standard and can produce as many as 750 output tokens per second.

OpenAI launches Ultrafast tier—GPT‑5.6 Sol up to 14× faster with Cerebras

OpenAI has introduced a new service tier called Ultrafast that runs GPT‑5.6 Sol up to 14× faster than its Standard processing mode. The feature is launching first in the OpenAI API and, using Cerebras hardware, can generate up to 750 output tokens per second.

Purpose and positioning

According to OpenAI, Ultrafast aims to make speed a competitive advantage by delivering higher throughput without sacrificing model intelligence. Previously, achieving real‑time responsiveness often required selecting smaller or more specialized models; Ultrafast is presented as a way to increase useful work per second while keeping the frontier capabilities of GPT‑5.6.

Early use cases and testing

OpenAI is running Ultrafast in a limited preview with an initial group of customers to identify where the increased speed provides the most value. Early tests involve companies across coding, commerce, financial research, support, and other interactive applications. By focusing first on business workflows, OpenAI intends to study the effects in production environments and learn how products and user interactions change when the model can keep pace with people.

Examples from internal usage

OpenAI’s internal developer teams have also trialed Ultrafast. One highlighted example is incident response: when an alert fires, engineers must form an accurate picture while systems and evidence are still changing. With Ultrafast, teams can more rapidly read logs, analyze traces, synthesize conversations, identify next checks, and prepare or validate fixes—shortening the time between observing a signal, testing hypotheses, and choosing actions. Human engineers remain responsible for judgment and deployment.

In research workflows, Ultrafast is used to quickly search knowledge sources, query data, and gather, organize, and summarize information across connected tools. Where teams previously launched batches of experiments overnight and reviewed results in the morning, Ultrafast can tighten that loop to enable multiple iterations during the same workday.

Partnership with Cerebras

Ultrafast continues OpenAI’s collaboration with Cerebras to deliver ultra‑low‑latency inference on OpenAI’s platform. Cerebras hardware supports GPT‑5.6 Sol in Ultrafast mode, enabling the model to produce up to 750 output tokens per second—intended to help businesses build more responsive products and incorporate advanced AI into time‑sensitive workflows.

Availability and next steps

GPT‑5.6 Sol on Ultrafast mode is available today in a limited preview to a select group of customers. OpenAI says it will expand access as capacity grows and invites businesses that need frontier intelligence at the highest speed to sign up for notifications when access widens.

Conclusion

Ultrafast represents a move by OpenAI to combine frontier model capabilities with much faster inference, targeting scenarios where latency and throughput materially affect outcomes. Initial customer and internal tests suggest particular benefits for incident response, research, and interactive applications; broader availability will depend on increasing capacity.