Business

OpenAI cuts GPT-5.6 Luna price by 80% and Terra by 20%, adds Fast mode for Sol

OpenAI sharply lowered prices for two models in its GPT-5.6 frontier series: Luna’s combined input-plus-output cost fell 80% to $1.40 per million tokens, and Terra’s fell 20% to $14 per million.

OpenAI cuts GPT-5.6 Luna price by 80% and Terra by 20%, adds Fast mode for Sol

OpenAI has significantly reduced prices for two models in its GPT-5.6 frontier family: Luna, the smallest and fastest model, received an 80% price cut, while Terra, the mid-tier model, was lowered by 20%. The company also added a premium Fast mode to its flagship Sol model.

OpenAI co-founder and CEO Sam Altman announced the changes on X, describing them as "major price cuts today."

New prices (per 1 million tokens)

  • GPT-5.6 Luna: $0.20 input, $1.20 output — $1.40 total.
  • GPT-5.6 Terra: $2.00 input, $12.00 output — $14.00 total.
  • GPT-5.6 Sol — Standard: unchanged at $5.00 input and $30.00 output — $35.00 total.
  • GPT-5.6 Sol — Fast: new premium at twice the Standard price: $10.00 input and $60.00 output — $70.00 total.

OpenAI says Fast mode can deliver up to 2.5x throughput without changing the model’s underlying intelligence.

Market context and impact

Luna’s 80% reduction moves its combined cost from $7.00 per million tokens (previously $1 input + $6 output) down to $1.40, positioning it much closer to the market’s lowest-cost commercial models. That new price undercuts Google’s Gemini 3.5 Flash-Lite ($2.80 total) and is far below Gemini 3.6 Flash ($9.00 total). Still, some vendors such as Xiaomi (MiMo-V2.5 Flash) and DeepSeek offer cheaper pure token pricing.

Terra’s 20% cut lowers its combined price from $17.50 to $14.00 per million tokens, now matching Google’s Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less and undercutting OpenAI’s GPT-5.4 ($17.50 total).

By contrast, Sol Fast’s $70.00 per million token rate makes it the most expensive configuration in the comparison, reflecting OpenAI’s choice to charge a latency premium for performance-sensitive workloads rather than lowering Sol’s base price.

Rival moves and different approaches to cost reduction

The timing follows recent competitive moves: Anthropic released Claude Opus 5 at the same combined $30 per million token rate as Opus 4.8, and Google introduced lower-cost models Gemini 3.6 Flash and Gemini 3.5 Flash-Lite focused on reducing inference costs and agent workload overhead.

Each vendor’s strategy differs:

  • OpenAI reduced per-token fees directly for Luna and Terra.
  • Google combines lower sticker prices with architecture and behavior changes that reduce token use and tool calls, lowering total runtime cost.
  • Anthropic delivered a more capable model at the same price as its predecessor, effectively lowering cost per unit of capability and added adjustable effort settings to trade reasoning depth for speed.

Third-party analyses cited in industry commentary indicate OpenAI’s GPT-5.6 models remain highly competitive on price-per-intelligence, with some benchmarks finding Luna outperforms Google’s Gemini 3.6 Flash and older Gemini models. Anthropic’s Claude Opus 5 remains competitive with Sol on performance while being about 6% cheaper on the advertised token rates.

Positioning of GPT-5.6 models

OpenAI describes the GPT-5.6 lineup as a frontier family released initially in late June 2026 via a limited rollout at U.S. government request before broader access. The three models are positioned with different trade-offs:

  • Sol: for the most complex, reasoning-heavy and agentic workloads, such as advanced coding and multi-step planning.
  • Terra: for general production use where capability and efficiency must be balanced.
  • Luna: for high-throughput, low-latency tasks (summarization, classification, routing, lightweight real-time assistants) where cost per request is the primary constraint.

Why this matters for deployments

For high-volume applications, even modest differences in token pricing can compound significantly across coding agents, document systems, internal search, and automated workflows. Luna’s 80% cut therefore materially alters OpenAI’s competitiveness in the low-cost inference tier, moving the company’s frontier-series model into direct competition with low-cost offerings from Google, Xiaomi, DeepSeek and others.

OpenAI’s move looks like a strategic repositioning rather than a routine price tweak: Sol remains the premium option, Terra shifts closer to pro-tier competitors, and Luna becomes OpenAI’s direct answer to the market’s growing low-cost segment. The broader question for enterprises is not just which model performs best, but which approach—lower per-token fees, fewer tokens used, or higher capability at the same price—yields the lowest total cost of production work.