OpenAI has made the GPT‑5.6 model family generally available after a limited preview. The family includes three tiers: GPT‑5.6 Sol (the flagship), GPT‑5.6 Terra (a balanced model for everyday work) and GPT‑5.6 Luna (the most cost‑efficient). OpenAI says the release focuses on getting more useful work from every token, improving performance per dollar, and offering capability scaling for demanding tasks.
Capabilities and how improvements are measured
According to OpenAI, GPT‑5.6 Sol sets a new standard in efficiency and intelligence: it achieves better results in coding, knowledge work, cybersecurity and scientific tasks while using fewer tokens and lower estimated cost. The company also introduced an “ultra” mode that coordinates multiple parallel agents to accelerate complex tasks.
On Agents’ Last Exam (which evaluates long‑horizon, agentic workflows across 55 professional fields), GPT‑5.6 Sol scores 53.6, 13.1 points higher than Claude Fable 5 (adaptive reasoning). At medium reasoning, Sol beats Fable 5 by 11.4 points at roughly one‑quarter of the estimated cost. OpenAI says these efficiency gains extend to the smaller models: Terra and Luna outperform Fable 5 at about one‑sixteenth of the cost.
On the Artificial Analysis Intelligence Index, GPT‑5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost.
Coding, terminal workflows and tool use
OpenAI describes GPT‑5.6 Sol as its best coding model to date. On the Artificial Analysis Coding Agent Index, Sol with max reasoning scores 80 — 2.8 points above Fable 5 — while using less than half the output tokens, taking less than half the time, and costing about one‑third less. Terra performs slightly above Fable 5, and Luna outperforms Opus 4.8; both do so with reduced time, fewer output tokens and lower estimated cost.
Sol also sets new state‑of‑the‑art results on Terminal‑Bench 2.1 and DeepSWE, which measure complex command‑line workflows and long‑horizon engineering in real codebases.
Technically, GPT‑5.6 can write and run lightweight programs that coordinate tools, process intermediate results, monitor progress and choose next actions as work unfolds. The Responses API’s Programmatic Tool Calling filters large amounts of intermediate data, retains only relevant pieces, and adapts workflows, which reduces token usage and model round trips.
More time and more agents: max, xhigh and ultra
The model offers multiple reasoning settings. For problems that benefit from greater compute and time, max gives GPT‑5.6 extra reasoning budget. Ultra goes further: by default it coordinates four parallel agents, trading higher token usage for stronger and faster results on demanding tasks. OpenAI’s comparisons indicate that adding parallel agents shifts the score‑latency frontier upward and left across several benchmarks.
Developers can build ultra‑like experiences using the multi‑agent beta in the Responses API.
Knowledge work, documents and presentations
GPT‑5.6 converts natural‑language requests into polished, interactive explanations and visualizations inside ChatGPT Work. It takes messy context from documents and collaboration tools (Slack, Notion, Microsoft 365, Google Drive) and produces shareable, expert‑level artifacts.
On BrowseComp, GPT‑5.6 Sol achieves a new state of the art at 92.2%. On OSWorld 2.0 Sol scores 62.6%; there it surpasses Opus 4.8 while using 85% fewer output tokens. Across the family, Luna nearly matches GPT‑5.5’s peak performance at less than half the estimated cost, while Terra exceeds GPT‑5.5 at a lower cost.
The company reports improvements in presentations, documents and spreadsheets: GPT‑5.6 can create fully editable presentations from scratch, infer and consistently apply a deck’s design system (layouts, typography, spacing, colors and Slide Master rules), and follow complex reference formats more faithfully. It also handles equations and financial models with greater precision.
Cybersecurity and defensive access
OpenAI says GPT‑5.6 is its strongest cybersecurity model so far while noting the dual‑use nature of these capabilities. On ExploitBench 1 (progress from discovering vulnerable code to arbitrary code execution), GPT‑5.6 scores 73.5% versus GPT‑5.5’s 47.9% at a comparable output‑token budget. On ExploitGym 2, the two‑hour pass rate rises from 15.1% (GPT‑5.5) to 24.9% (GPT‑5.6); at six hours the pass rate reaches 33.7%. On SEC‑Bench Pro (proof‑of‑concept generation on complex software), GPT‑5.6 scores 71.2% versus GPT‑5.5’s 45.8% with improved latency.
To preserve defensive uses while limiting misuse, OpenAI offers Trusted Access for Cyber through the Daybreak program: qualified individuals and organizations can gain access to more precise defensive capabilities in authorized environments. Individuals can verify identity and request Trusted Access; organizations can apply for team access. Members who want retained access to the company’s most cyber‑capable frontier models must enable Advanced Account Security.
Scientific research and internal adoption
OpenAI reports Pareto improvements over GPT‑5.5 on life‑sciences and chemistry evaluations and stronger results on long‑horizon genomics and quantitative‑biology workflows (GeneBench Pro). Internally, researchers used GPT‑5.6 across the development loop: diagnosing failures, optimizing training systems, running experiments and interpreting results. During internal testing, average daily output tokens per active researcher exceeded the highest level seen for GPT‑5.5 by more than twofold.
An internal evaluation bundle measuring progress toward recursive self‑improvement (RSI) showed GPT‑5.6 Sol improved by 16.2 points over GPT‑5.5.
Safety, red‑teaming and monitoring
OpenAI says it deployed its most extensive safety evaluations to date for GPT‑5.6, combining human red teaming, large‑scale automated testing and external expert review. They report roughly 700,000 A100e GPU hours of black‑box automated red teaming prior to general availability. The safety stack combines protections trained into the model with real‑time checks, continuous monitoring, access controls and a reasoning monitor that reviews conversations for potential harm.
OpenAI also reports that GPT‑5.6 Sol cyber safeguards block approximately ten times more potentially harmful activity than previous models. Because such protections can impede benign uses, the company provides options in ChatGPT and Codex to retry prompts on lower‑capability models and says it will iteratively reduce the friction of safeguards while maintaining robustness.
OpenAI acknowledges no security is perfect: new weaknesses and jailbreaks will appear, so they pair existing security and biology bug bounty programs with a rapid‑remediation process and increased monitoring. Findings from researchers and real‑world misuse will feed into ongoing evaluations and safeguards.
Availability, pricing and technical changes
GPT‑5.6 is available immediately across ChatGPT, Codex and the OpenAI API; the global rollout has started and will proceed toward full availability over the next 24 hours.
Pricing per 1M tokens:
- GPT‑5.6 Sol: $5 input / $30 output
- GPT‑5.6 Terra: $2.50 input / $15 output
- GPT‑5.6 Luna: $1 input / $6 output
GPT‑5.6 introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30‑minute minimum cache life. Cache writes are billed at 1.25× the model’s uncached input rate; cache reads keep a 90% cached‑input discount.
Conclusion
OpenAI positions the GPT‑5.6 family as an efficiency‑focused step forward: higher task performance per token and per dollar, new multi‑agent operation modes for demanding work, and strengthened safeguards developed through extensive red‑teaming and automated testing. The models are available now with tiered pricing and Trusted Access pathways for higher‑risk defensive capabilities.



