OpenAI announced price cuts for two members of the GPT‑5.6 family and introduced a new Fast mode for GPT‑5.6 Sol in the API. The company says the changes — effective in the API from July 30 — reflect efficiency improvements across model design, serving systems and agent tooling, and are intended to let customers get more work per dollar and move faster when latency matters.
What changes for customers
- Immediate price reductions: GPT‑5.6 Luna, described as the fastest and most affordable model, will be priced 80% lower; GPT‑5.6 Terra, the balanced model for everyday work, will be priced 20% lower.
- API token pricing (effective July 30):
- Terra: $2 per million input tokens and $12 per million output tokens.
- Luna: $0.20 per million input tokens and $1.20 per million output tokens.
- Sol pricing remains unchanged.
- Fast mode in the API: replacing the previous Priority Processing, Fast mode for GPT‑5.6 Sol offers up to 2.5× faster speeds than Standard processing at twice the price. OpenAI says Fast mode does not change model intelligence and is backwards compatible: requests tagged as priority will automatically use Fast mode.
Expected practical effects
OpenAI argues that the Luna price cut makes high‑volume, high‑quality workloads more cost‑effective. The company states that Luna delivers performance comparable to models that were frontier‑class a year ago, at "roughly 6 cents on the dollar per task" and with nearly nine times the speed. On professional work measured by Agents’ Last Exam, OpenAI says Luna outperforms Fable 5 while costing an estimated nearly 99% less per task.
In practice, businesses can mix models across a workflow: for example, use Sol to resolve uncertainty and plan, then use Luna to implement well‑specified changes, write and run tests, and evaluate results. Different workflows can choose different balances of cost, speed and intelligence.
What underpins the savings
OpenAI credits multiple engineering improvements for the cost reductions: better model architectures, more efficient inference systems, and an agentic layer that connects models to tools and context. The company highlights concrete gains:
- GPT‑5.6 Sol autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training, intervening when necessary.
- Kernel work reduced end‑to‑end serving cost by about 20%.
- Token‑generation experiments increased token‑generation efficiency by more than 15%.
OpenAI frames these changes as a feedback loop: more capable models can autonomously discover efficiency improvements that then allow further gains.
Infrastructure and use cases
Meeting demand for abundant intelligence, OpenAI says, requires both more compute and more productive compute. The company pairs workloads with systems suited to them so the approach supports both low‑cost, high‑volume use (Luna and Terra) and frontier, latency‑sensitive workloads (Sol with Fast mode).
The firm points to large‑scale document analysis, customer interaction classification and routine implementation as examples of workloads that should become economical to run broadly, while reserving faster Sol processing for the tasks where the premium is justified.
Access and subscriptions
GPT‑5.6 Terra and Luna remain available in ChatGPT Work, Codex and the OpenAI API. In ChatGPT Work and Codex, Free and Go users can access Terra; Plus, Pro, Business and Enterprise subscribers can choose both Terra and Luna. ChatGPT and Codex subscription prices and quota budgets do not change, but Terra and Luna usage will consume fewer credits.
Rollout schedule
OpenAI says the API price changes take effect July 30 with the listed per‑token rates; pricing updates will begin rolling out on AWS later the same day. Fast mode replaces Priority Processing in the API and aligns with /fast in Codex; existing API requests tagged priority will continue to work.
Summary
OpenAI is passing model and serving efficiency gains to customers by lowering Luna and Terra prices and by offering a faster Sol processing mode in the API. The company positions these moves as enabling enterprises to optimize the tradeoffs among intelligence, speed and cost across multi‑step workflows while maintaining access to faster responses when it matters most.



