Google breaks the price of the 'workhorse' tier while open-weight LongCat-2.0 takes the lead on OpenRouter
On July 21, Google launches Gemini 3.6 Flash at $1.50 / $7.50 and a Flash-Lite at $0.30 / $2.50 per million tokens and teases Gemini 4; the same market sees open model LongCat-2.0 rise among OpenRouter's most-used and Meta open its first paid API at a quarter of the leaders' price. The mid-tier collapses while the frontier locks down.

Two days after Anthropic opened its IPO roadshow, the AI market shifted to where it really matters for an organization paying for its tokens: the mid-tier. On July 21, Google launched a trio of Gemini Flash models built for throughput and cost, and hinted at Gemini 4. The same day, China’s open model LongCat-2.0, trained without a single Nvidia GPU, already ranks among the most-consumed models on OpenRouter, and Meta, long the champion of free open-weight, started charging for its API for the first time. Three signals pointing the same way: the price of routine work is collapsing and opening up, while the frontier locks itself behind meters.
Google reopens the price war on the “workhorse” tier
On July 21, 2026, Google shipped three models at once: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, all positioned for efficiency, low latency and reliability to run agents at scale. The lineup’s workhorse, Gemini 3.6 Flash, is priced at $1.50 / $7.50 per million tokens (input / output), output drops below the $9.00 of the previous 3.5 Flash. Alongside, Flash-Lite comes down to $0.30 / $2.50 for very high throughput, while Flash Cyber, dedicated to detecting and patching vulnerabilities via Google’s CodeMender agent, remains in a restricted pilot reserved for governments and trusted partners over dual-use concerns.
The key isn’t just the sticker price. Google emphasizes a gain in efficiency per task: 3.6 Flash reportedly consumes about 17% fewer output tokens than 3.5 Flash for the same work, with fewer reasoning steps and tool calls across multi-step chains, which, on top of the lower rate, pulls the effective cost of a task down well beyond what the sticker says (up to a claimed ~71% on agentic coding workloads). Its knowledge cutoff finally advances from January 2025 to March 2026. In the background, Google teased Gemini 4 with no firm timeline, while its 3.5 Pro promised for June remains, for its part, absent. The message to the market is clear: the battle is now won on cost per completed task, not merely cost per million tokens.
LongCat-2.0: open-weight settles at the top of OpenRouter
The second tremor comes from China, and it is structural. LongCat-2.0, released by Meituan on June 30, 2026 under an MIT license, is a Mixture-of-Experts model of 1.6 trillion parameters (≈ 48 billion active per token) with a native one-million-token context. Two facts make it notable beyond its benchmarks (SWE-bench Pro at 59.5, ahead of GPT-5.5’s 58.6 per the vendor’s figures). First, after running anonymously under the name “Owl Alpha” for nearly two months, it climbed among OpenRouter’s most-used models, with a reported volume of roughly 10.1 trillion tokens per month and a +242% month-over-month jump; VentureBeat even describes it as “leading” the platform. Second, Meituan claims to have trained it entirely on more than 50,000 Chinese-made chips (ASICs), with no Nvidia GPUs, the first trillion-parameter model built on domestic compute.
On price, the gulf with the Western frontier is glaring: $0.75 / $2.95 per million tokens at standard rates, and $0.30 / $1.20 on promotion, with cache hits billed free. A near-frontier model, open, self-hostable, at one-twentieth of a Claude Fable 5’s output price: this is precisely the argument China staged in Shanghai last week, and it is no longer rhetoric but market share.
Meta starts charging for the first time
The third signal, quieter but heavy with meaning: Meta, whose default doctrine was free open-weight, opened a paid API on July 9 for its Muse Spark 1.1 model, priced at $1.25 / $4.25 per million tokens ($20 in credits at signup, a cache rate floated around $0.15 but unofficial). Mark Zuckerberg presents it as “the first time that we’re doing a real serious API,” with pricing that is “very aggressive and attractive”, in effect, about a quarter of what OpenAI or Anthropic charge for comparable models. The API natively accepts OpenAI (Chat Completions) and Anthropic (Messages) formats, to allow switching with no migration cost. In other words: even the historically most open player is now commercializing access, but by breaking prices from below.
Comparison table: API prices as of July 22, 2026
The landscape remains stretched across a factor of one hundred between the priciest and the most economical, but the bottom and middle of the table have thickened in a single week.
| Model | Input ($/M tokens) | Output ($/M tokens) | Note |
|---|---|---|---|
| Claude Fable 5 (Anthropic) | 10.00 | 50.00 | frontier; subscription limits tightened July 20 |
| GPT-5.6 Sol (OpenAI) | 5.00 | 30.00 | launched July 9 |
| Claude Opus 4.8 | 5.00 | 25.00 | |
| GPT-5.6 Terra | 2.50 | 15.00 | launched July 9 |
| Kimi K3 (Moonshot, China, open-weight) | 3.00 | 15.00 | launched July 16; cache-hit 0.30 |
| Gemini 3.1 Pro (Google) | 2.00 | 12.00 | doubles beyond 200k context |
| Claude Sonnet 5 | 2.00 | 10.00 | launch price until August 31, then 3/15 |
| Gemini 3.6 Flash (Google) | 1.50 | 7.50 | launched July 21; −17% tokens/task |
| Grok 4.5 (xAI) | 2.00 | 6.00 | price published; EU access still restricted |
| GPT-5.6 Luna | 1.00 | 6.00 | launched July 9 |
| Meta Muse Spark 1.1 | 1.25 | 4.25 | launched July 9; Meta’s first paid API |
| LongCat-2.0 (Meituan, open-weight) | 0.75 | 2.95 | promo 0.30/1.20; MIT; leading OpenRouter |
| Gemini 3.5 Flash-Lite (Google) | 0.30 | 2.50 | launched July 21; very high throughput |
| Grok 4.3 / 4.20 (xAI) | 1.25 | 2.50 | 1M context |
| Mistral Large 3 | 0.50 | 1.50 | |
| DeepSeek V4 Pro | 0.435 | 0.87 | after the May cut |
| DeepSeek V4 Flash | 0.14 | 0.28 | 1M context; cache up to −99% |
Prices collected on July 22, 2026 from official pages (Anthropic, OpenAI, Google, Meta) and specialized aggregators for models without a directly accessible pricing page. Promotional and cache rates are flagged in the note column; they can expire without notice.
The frontier, for its part, locks down, and compute stays the blind spot
While the mid-tier goes on sale, the top of the table takes the opposite path. Claude Fable 5 stays at $10 / $50 and, since July 20, its Max and Team Premium subscription limits have been cut by roughly a third, with Pro and Team Standard users shifting to metered billing. Anthropic’s IPO roadshow, listing targeted as early as October, opening valuation that prediction markets place around $1.1 trillion, assumes tight monetization of the most capable model, not an unlimited plan. The market is splitting into a barbell: a rare, metered frontier on one side, an open and cheap abundance on the other.
Beneath this shift, the hardware bedrock isn’t budging, but it is recomposing. Renting an H100 still runs between ~$2 and ~$11 an hour depending on the provider, and a purchased server’s break-even still demands long months of sustained use. Two deeper moves, however, change the equation: LongCat-2.0 shows that a cutting-edge model can be trained without American silicon, which shifts the balance of power over compute supply; and consolidation is accelerating, Qualcomm announced the acquisition of AI platform Modular for roughly $3.9 billion to strengthen its infrastructure from the edge to the datacenter. The token price is falling; mastering the compute that carries it is becoming a geopolitical stake.
What this means for a Swiss organization
Three concrete points of attention. First, the real cost lever is no longer the per-token rate but the number of tokens consumed per task: the Gemini Flash update shows that a cheaper and more output-frugal model can cut the actual bill well beyond its posted price drop, provided you measure your own consumption to see it. Second, open-weight is no longer a fallback alternative but a distribution standard: with LongCat-2.0 among OpenRouter’s most-used models and an MIT license, sovereign self-hosting becomes a governance and cost choice with no performance compromise. Finally, the market’s barbell structure calls for arbitrage, not picking a side: reserve the metered frontier for the tasks that demand it, route routine volume toward a mid-tier that has turned abundant and cheap, and keep control of that routing, since limits and rates can change overnight, at the top as at the bottom.
Control your AI costs instead of enduring them. Comoto OS orchestrates the market's best models, proprietary or open-weight, from a sovereign core hosted in Switzerland: every task routed to the right model at the right price, every token traced, every decision auditable, whatever limits and rates your provider decides on tomorrow.
Discover Comoto OSBe the first to read our upcoming papers.
Some articles are reserved. Leave your email to unlock access to upcoming Meotis Research publications.




