NewsroomNews · AI Market

GPT-5.6 and Grok 4.5 launch the same day, and the cost war intensifies

A "Super Thursday" of frontier launches, API prices spread across a factor of 100, and inference now devouring more than half of infrastructure spending. A tour of the numbers.

Black and white engraving: a rock with a glowing core fractures into fragments descending like a price curve

July 9, 2026 will be remembered as a rare day: OpenAI publicly launches its GPT-5.6 series on the very day xAI releases Grok 4.5. Behind these simultaneous announcements lies a more structural reality for any organization that consumes AI: the price gap between frontier models now stretches across a factor of 100, and knowing how to arbitrate between them has become a financial skill in its own right.

A “Super Thursday” of frontier launches

OpenAI had confirmed the day before that its GPT-5.6 series, three models named Sol, Terra, and Luna, would arrive that Thursday on ChatGPT, the API, and Codex, after several weeks of preview reserved for partners. The prices reported at launch: Sol at $5/$30 per million tokens (input/output), Terra at $2.50/$15, and Luna at $1/$6. A telling detail flagged by early benchmarks: Luna, the cheapest model in the series, reportedly beats Terra on some agentic test suites.

That same day, xAI opened access to Grok 4.5, presented as a frontier-class model, for now reserved for SuperGrok Heavy and Premium+ subscribers, with no published API price and no independent benchmarks at the time of writing. Two major launches in a single day: the sign of a market where release cadence has become a competitive weapon in itself.

The price war, reignited by Asia

The pressure on prices is not coming from the American giants. In May, DeepSeek cut its rates by about 75%: its V4 Pro now costs $0.435/$0.87 per million tokens, and V4 Flash drops to $0.14/$0.28 with a one-million-token context window. According to figures cited by CNBC in early July, 30 to 46% of the tokens consumed by American companies already run through Chinese models, billed 60 to 90% less than their Western equivalents.

At the other end of the spectrum, Anthropic has turned its most capable model into an unapologetically premium product: since July 8, Claude Fable 5 has been billed at $10/$50 per million tokens, accessible to subscribers through a credit system. Between the most expensive and the most economical model on the market, the price gap reaches a factor of 100 on input: unprecedented.

The price of a token is no longer a technical detail buried in a cloud invoice. It has become a strategic parameter, on par with the choice of an energy supplier.

Comparison table: API prices as of July 9, 2026

Model Input ($/M tokens) Output ($/M tokens) Note
Claude Fable 5 (Anthropic) 10.00 50.00 credits add-on since July 8
GPT-5.6 Sol (OpenAI) 5.00 30.00 launched July 9
Claude Opus 4.8 5.00 25.00
GPT-5.6 Terra 2.50 15.00 launched July 9
Gemini 3.1 Pro (Google) 2.00 12.00 doubles beyond 200k context
Claude Sonnet 5 2.00 10.00 launch price until August 31, then 3/15
Gemini 3.5 Flash 1.50 9.00
Grok 4.3 (xAI) 1.25 2.50 1M context
GPT-5.6 Luna 1.00 6.00 launched July 9
DeepSeek V4 Pro 0.435 0.87 after the −75% cut (May)
Grok 4.1 Fast 0.20 0.50
DeepSeek V4 Flash 0.14 0.28 1M context

Prices collected on July 9, 2026 from specialized aggregators and official pages. GPT-5.6 rates are from launch day, and Grok 4.5 does not yet have a published API price: confirm on the vendors’ official pages.

On the subscription side: the announced end of unlimited

Consumers are not spared either. ChatGPT now ranges from the Go plan at $8 up to Pro at $200 per month, with an intermediate tier at $100 launched in April to answer Anthropic’s Claude Max. At Anthropic, meanwhile, Pro remains at $20 and Max tops out at $200, but access to the frontier model Fable 5 now goes through credits billed at API prices, including for subscribers. An OpenAI executive even publicly compared unlimited plans to “unlimited electricity”: unsustainable in the long run, with a Plus subscription that could reach $44 by 2029.

The message is consistent from one vendor to the next: heavy AI usage will be paid for by consumption, and organizations have every interest in measuring theirs precisely.

Why prices are falling on the surface and rising behind the scenes

An apparent paradox: per-token prices are falling, yet spending is exploding. Inference, running the models rather than training them, now accounts for 55% of AI infrastructure spending, up from a third in 2023. H100 GPU rental rates have regained about 40% in five months, driven by that demand. And hyperscalers are planning $400 to $750 billion in datacenter capex for 2026 depending on the scope: Anthropic, for instance, just signed a $19 billion lease with TeraWulf.

On the funding side, the first half of 2026 broke every record: $510 billion invested in venture capital, 43% of it captured by OpenAI and Anthropic alone. According to Fortune, Anthropic may even have overtaken OpenAI in annualized revenue (~$47 billion self-reported, unaudited figures). Value is concentrating, costs are shifting toward usage, and the final bill lands on the companies that use these models.

What this means for a Swiss organization

Three practical consequences. First, multi-model is no longer optional: routing each task to the model with the right capability-to-price ratio can cut the bill by a factor of ten with no perceptible loss of quality. Second, cost traceability is becoming a governance issue: without fine-grained measurement of who consumes what, arbitration is impossible. Finally, dependence on a single provider has become a pricing risk: price changes are now decided in days, not years.

Comoto OS · Sovereign cognitive core

Control your AI costs instead of enduring them. Comoto OS orchestrates the market's best models from a sovereign core hosted in Switzerland: every task routed to the right model, every token traced, every decision auditable.

Discover Comoto OS
Early access

Be the first to read our upcoming papers.

Some articles are reserved. Leave your email to unlock access to upcoming Meotis Research publications.