OpenAI introduced GPT‑6 Sol and GPT‑6 Luna, stating that the new models' API list prices are about 50% lower than the promotional rates for GPT‑5.6 Sol and Luna. At nearly the same time, Anthropic released Claude Opus 5.5 and said that, under typical loads, the total running cost of Opus 5.5 can be roughly 40% lower than Opus 5.
At first glance these are technical price changes, but in enterprise settings they have a direct impact: the operating cost of continuously running AI agents is determined principally by per‑token pricing for input and output.
The numbers from list pricing
According to the standard API list prices cited in the sources, the changes are as follows (USD per 1 million tokens):
- GPT‑5.6 Sol → GPT‑6 Sol: input 4.0 → 2.0 (-50%), output 20.0 → 10.0 (-50%)
- GPT‑5.6 Luna → GPT‑6 Luna: input 0.2 → 0.1 (-50%), output 1.2 → 0.5 (-58%)
- Claude Opus 5 → Claude Opus 5.5: input 5.0 → 4.0 (-20%), output 25.0 → 20.0 (-20%)
These are list prices; actual enterprise bills depend on prompt caching, task length, processing mode and model token efficiency.
Practical impact
In API‑based usage companies often execute thousands or tens of thousands of calls per day: each run involves the system “reading” prompts, background documents or prior conversation as input tokens, then the model producing answers, summaries or code as output tokens. In a simplified example, processing 1 million input and 1 million output tokens would have cost $24 with GPT‑5.6 Sol but $12 with GPT‑6 Sol; for the Luna line the cost falls from $1.40 to $0.60 in the same scenario.
At enterprise scale, where such calls run continuously, those savings translate into meaningful operational cost reductions.
Cost‑efficiency as a competitive metric
The market is shifting from pure benchmark competition toward cost‑efficiency: vendors must show not only which model scores higher on tests, but how much useful work a model delivers per dollar. OpenAI highlighted improvements to prompt caching, which can lower costs for repeated or reused context in long conversations and agent workflows. Anthropic says Opus 5.5 is cheaper per token, uses fewer tokens for some tasks, made cached content rereads less expensive, and offers over 30% faster response generation in its claims.
Security and geopolitical considerations
Lower prices increase accessibility, which raises questions about who can run high‑capability models and under what controls. Anthropic noted that Opus 5.5 underwent external evaluations and that it applies advanced safeguards similar to its most sophisticated systems, with stricter handling for biological and cybersecurity use cases and improved resistance to prompt injection attacks.
There is also a geopolitical aspect: reporting summarized by the AP indicates that Chinese models led by Moonshot AI and DeepSeek can already challenge American models on several measures and often at lower cost. That does not prove U.S. price cuts were driven solely by Chinese competition, but timing and market context suggest the pressure is relevant.
Conclusion
The combination of list price reductions and technical improvements (token efficiency, better caching, faster generation) can materially lower the cost of running enterprise AI agents. That may accelerate adoption of automated assistants across business processes, while keeping security, controls and regulation on the agenda as critical counterweights.
Note: the source mentions a Portfolio AI & Digital Transformation conference scheduled for 26 November 2026; this is ancillary promotional information included in the original material.



