NewsroomNews · AI Market

The top of the global leaderboard just halved in price, and the largest open model in history ships today

Claude Opus 5 edges past Fable 5 on the Artificial Analysis index while costing half as much, and Moonshot releases Kimi K3's 2.8 trillion parameters on July 27. The question is no longer which model is best, but which stays cheapest at equal capability.

Black and white engraving: a monumental stone balance scale where a small crystalline geode with a blazing core, half the size of its counterweight, lifts an enormous cracked monolith off its pan, value no longer measured by mass

For two years the AI market ran on a simple rule: the most intelligent model was the most expensive one, and you paid the difference. That rule just broke. On July 24, Anthropic launched Claude Opus 5, which takes first place on the Artificial Analysis Intelligence Index with 61 points, ahead of Claude Fable 5 (60) and GPT-5.6 Sol (59), while being billed at $5 / $25 per million tokens, exactly half the rate of Fable 5. And on July 27 at 00:00 UTC, Moonshot releases the weights of Kimi K3, 2.8 trillion parameters, the largest open model release in history. The top of the market is coming down, the bottom is becoming free. For an organization paying for its tokens, the question of the year is no longer “which model is best,” but “what does a completed task actually cost me.”

Opus 5 takes the lead while costing half of Fable 5

The move is unusual: a vendor dethrones its own flagship with a model that costs twice as little. Claude Opus 5 is billed at $5 input and $25 output per million tokens, the exact rate of its predecessor Opus 4.8, for a level of performance that now places it ahead of Fable 5 ($10 / $50). Anthropic frames the proposition plainly: Opus 5 “provides greatly improved performance for the same cost as its predecessor, Opus 4.8.”

The number that matters to a finance director, however, is not the per-token rate but the cost per completed task, since a more efficient model burns fewer reasoning tokens to reach the end. On that ground the gap is clear. Artificial Analysis measures a weighted average cost of $2.03 per Intelligence Index task for Opus 5 (max effort), against $2.75 for Fable 5, or 26% less at comparable intelligence. On its proprietary agentic knowledge-work benchmark, AA-Briefcase, the gap widens: Opus 5 reaches an Elo of 1,720 against 1,574 for Fable 5, a 146-point lead, for a cost per task of $17.79 against $22.30 (−20%). The intermediate configurations are more favorable still: the xhigh and high effort levels both beat Fable 5 while costing only 64% ($14.26) and 47% ($10.41) of its price respectively.

On raw capability, the model posts 96.0% on SWE-bench Verified, effectively saturating the reference benchmark for software engineering, and 79.2% on the more demanding SWE-bench Pro. On access, the context window is 1 million tokens in and 128,000 out, extendable to 300,000 output tokens via the Batch API with a beta header. A Fast mode runs at roughly 2.5 times the default speed for twice the base price. Finally, a detail that carries weight in regulated environments: unlike Fable 5, Opus 5 carries no data retention requirement. On subscriptions, it becomes the default model on Claude Max and the most capable one available to Claude Pro subscribers ($20 per month, $17 billed annually; Max at $100 or $200).

The gains are not uniform, and that is the important part

It is worth resisting the triumphal reading. Artificial Analysis itself calls Opus 5 “narrowly the most intelligent” model, 61 against 60 being a near-tie rather than a generational leap. The real change in this release is not raw capability, it is control over effort, and therefore over spend.

Independent measurements confirm this, and they are nuanced. On code review tasks, Opus 5’s x-high configuration produces more precise actionable comments (39.3% vs 35.2%), but catches fewer known issues (55.2% vs 61.1%) and generates roughly four times more nitpicks. Counting the full post-pipeline stream, its precision even falls below baseline (28.6% vs 32.8%): the model is wordier. In other words, effort has become a routing decision: moving up a notch buys precision at the cost of coverage, and nothing improves uniformly. For an organization, this means a headline 26% cost gain only materializes if you pick the right effort level for the right task, which requires measuring rather than assuming.

Kimi K3 opens 2.8 trillion parameters, and reveals the real price of “free”

While the top comes down, the base opens up. Since July 27, 00:00 UTC, the weights of Kimi K3 are published on Hugging Face: 2.8 trillion parameters in a mixture-of-experts architecture, of which only 16 experts out of 896 fire per token (about 50 billion active parameters), with a 1 million token context window. It is the largest weight release ever made, and on paper it is free.

In practice, the bill shifts to hardware. The weight file runs to roughly 1.4 TB in MXFP4 and 5.6 TB at 16-bit, which implies on the order of eighteen 80 GB accelerators just to load the model; a single node of eight cards at 192 GB (about 1.5 TB) holds the weights “with almost nothing to spare.” Compatible hardware is in practice limited to Nvidia Blackwell or AMD MI400. No consumer card will do, even quantized. Two caveats deserve explicit mention: sources diverge on the exact size of the MXFP4 package, some reporting about 594 GB rather than 1.4 TB, and above all the license terms had not been published at the time of the announcement, leaving commercial use unconfirmed. To this add the context of July 22: Kimi K3 is the subject of an official distillation accusation brought by the White House against Moonshot, an accusation contested by several experts and to date not technically corroborated. Via the API, the model remains listed at $3 / $15 since July 16, and scores 57 on the Artificial Analysis index, second on some leaderboards and first in the frontend coding arena.

The lesson is useful to anyone comparing costs: an open model is not a free model. It shifts the spend from the token to hardware amortization, operations and compliance. For most organizations, open weights are worth having first as continuity insurance, the guarantee of being able to bring a workload back in-house if the vendor changes its prices, its limits or its availability, and not as an immediate saving.

Beneath access prices, compute still isn’t budging

The hardware bedrock keeps swelling while access rates fall, and that is the market’s central tension. Nvidia posted record datacenter revenue of $75.2 billion, up 92% year over year, including $60.4 billion for compute (+77%) and $14.8 billion for networking (+199%), with guidance of $91.0 billion (±2%) for the following quarter. On rental, an H100 trades between $1.38 and $1.49 an hour on specialist marketplaces and up to $11.68 to $12.29 an hour at hyperscaler on-demand rates, a factor of eight for the same chip. Multi-year commitments take 25 to 50% off the on-demand rate, and specialist clouds (CoreWeave, Lambda, Nebius) stay 40 to 70% below AWS, Azure and Google Cloud. To buy, the card runs $25,000 to $40,000, an eight-card DGX H100 node $300,000 to $460,000.

Financing follows the same trajectory. OpenAI filed a confidential IPO prospectus with the SEC around May 22, publicly confirmed on June 8, with Goldman Sachs, Morgan Stanley and JPMorgan, targeting a September 2026 debut at a valuation of $730 to $850 billion, with some analysts expecting north of a trillion. The company reports $2 billion in monthly revenue, of which more than 40% from enterprise, but would still be losing about $1.22 for every dollar earned over the quarter, to be set against its compute envelope of $750 billion through 2030. In Europe, the Financial Times reported on July 22 that Samsung is negotiating up to €1 billion in Mistral at a valuation of roughly €20 billion, within a larger round involving EQT, Novo Holdings and Santander. xAI is reportedly discussing a further $10 billion at a $75 billion valuation, and Thinking Machines Lab targeting $1 billion at $9 billion (round not finalized). These last three figures are reported discussions, not closed transactions.

Comparison table: API prices as of July 27, 2026

The week’s only pricing move is the arrival of Opus 5, which replaces Opus 4.8 at the same rate but overturns the hierarchy at the top of the table: the first line of the index is no longer the most expensive one.

Model Input ($/M tokens) Output ($/M tokens) Note
Claude Fable 5 (Anthropic) 10.00 50.00 dethroned on the AA index July 24; data retention required
GPT-5.6 Sol (OpenAI) 5.00 30.00 AA index 59
Claude Opus 5 (Anthropic) 5.00 25.00 launched July 24; #1 on the AA index (61); $2.03/task; Fast mode ×2.5 at twice the price
GPT-5.6 Terra (OpenAI) 2.50 15.00 launched July 9
Kimi K3 (Moonshot, China, open-weight) 3.00 15.00 weights opened July 27 (1.4 TB); license unpublished; AA index 57
Gemini 3.1 Pro (Google) 2.00 12.00 doubles beyond 200k context
Claude Sonnet 5 2.00 10.00 launch price until August 31, then 3/15
Gemini 3.6 Flash (Google) 1.50 7.50 launched July 21; −17% tokens/task
Grok 4.5 (xAI) 2.00 6.00 coding model; EU access still restricted
GPT-5.6 Luna (OpenAI) 1.00 6.00 launched July 9
Inkling (Thinking Machines, open-weight) 1.87 4.68 Apache 2.0 license; multimodal
Meta Muse Spark 1.1 1.25 4.25 Meta’s first paid API
LongCat-2.0 (Meituan, open-weight) 0.75 2.95 promo 0.30/1.20; MIT license
Gemini 3.5 Flash-Lite (Google) 0.30 2.50 launched July 21; very high throughput
Grok 4.3 / 4.20 (xAI) 1.25 2.50 1M context
Mistral Large 3 0.50 1.50 Samsung in talks to invest (FT, July 22)
DeepSeek V4 Pro 0.435 0.87 deepseek-reasoner alias removed July 24
DeepSeek V4 Flash 0.14 0.28 1M context; deepseek-chat alias removed July 24

Opus 5, Fable 5, GPT-5.6 Sol and Kimi K3 rates verified on July 27, 2026 against official pages and Artificial Analysis. The other lines are carried over from the July 24, 2026 collection, no pricing change having been published since. Promotional and cache rates can expire without notice. Reminder: since July 24 at 15:59 UTC, the deepseek-chat and deepseek-reasoner identifiers no longer work, migrate to deepseek-v4-flash and deepseek-v4-pro.

What this means for a Swiss organization

Three operational conclusions. First, stop comparing per-token prices and start comparing costs per task. The gap between $2.03 and $2.75 per task at equivalent intelligence appears on no pricing page: it only shows up by measuring the real consumption of your own workloads. An organization that does not measure its cost per task cannot capture the decline, it will keep paying the sticker rate of the most prestigious model.

Second, the effort level has become an economic decision. Opus 5 demonstrates that at constant model, the choice of configuration swings the cost by a factor of two, with gains that vary by task. Routing each task to the right model and the right effort level, then verifying the result on your own data rather than on a public leaderboard, is where the real saving sits.

Finally, open weights are insurance, not a saving. With 1.4 TB of parameters, an unpublished license and a distillation accusation pending, Kimi K3 illustrates that a model’s sovereignty is paid for in infrastructure and legal risk. The right posture is neither to bring everything back in-house nor to entrust everything to a single vendor: it is to keep the ability to change your mind, to move a workload from a proprietary model to an open one without rewriting your foundation, whatever the prices, licenses and leaderboards the market decides on next month.

Comoto OS · Sovereign cognitive core

Control your AI costs instead of enduring them. Comoto OS orchestrates the market's best models, proprietary or open-weight, from a sovereign core hosted in Switzerland: every task routed to the model and effort level with the right cost/capability ratio, every token traced and measured as a real cost per task, and the freedom to switch models without rewriting anything, whatever tomorrow's prices and leaderboards may be.

Discover Comoto OS
Early access

Be the first to read our upcoming papers.

Some articles are reserved. Leave your email to unlock access to upcoming Meotis Research publications.