Seoul drops 10.84% and corporate America shuts off the token tap: the AI bill is coming due
The Kospi posts its worst session since 2008 on chip names while an AP investigation documents the end of "tokenmaxxing": Uber burned through its annual AI budget in four months, Salesforce is heading for a $300 million Anthropic bill. The price of a token has never been lower, and the invoices have never been heavier.

For eighteen months the AI market could be read on a single line: the price of a million tokens, which would not stop falling. On July 28, two stories published hours apart moved the question elsewhere. In Seoul, the Kospi lost 10.84% in a single session, its worst since 2008, carried down by the abrupt revision of one assumption: do AI revenues actually justify the semiconductor spending committed for them? And in Washington, an AP investigation by Matt O’Brien documented the symmetrical reversal on the enterprise side: the end of “tokenmaxxing,” the fad of consuming as many tokens as possible, now caught up with by the invoices. Uber exhausted its 2026 AI budget in four months. Salesforce is heading for $300 million in Anthropic billing on the year. The unit price of a token has never been lower, and that is precisely what made total spending explode.
The Kospi posts its worst session since 2008
The session’s numbers leave no ambiguity. The Korean index shed 732.09 points to close at 6,023.66, or −10.84%, after falling as far as −11.3% intraday and dropping below 6,000 points for the first time since April 14. The Korea Exchange triggered a circuit breaker after a decline of more than 8%, suspending trading for twenty minutes — the eighth activation this year. Of the 917 stocks on the main board, 878 fell and only 36 advanced.
The index’s two heavyweights concentrated the shock: Samsung Electronics fell 14.4%, its steepest daily drop since October 2008, and SK hynix 14.7%, its American depositary receipts slipping back below their offering price. Contagion was regional without being uniform: Nikkei 225 at −4.3%, Taiex at −4.7%, Shanghai Composite at −1.4%, while the Philadelphia Semiconductor index gave up 2.2% and Nvidia around 5%. Across the month of July, the Kospi is down 29%, a monthly decline exceeding the −27% record of October 1997, at the height of the Asian financial crisis.
Two triggers are cited, and they are of different orders. The first is industrial: Chinese memory maker ChangXin Memory Technologies surged 466% on its Shanghai debut, against a backdrop of reports of state-backed domestic lithography equipment production — that is, the prospect of competition in the exact segment that gives Samsung and SK hynix their value. The second is financial: the broad reappraisal of the relationship between AI infrastructure spending and observed revenue. One indication on that front: in a recent report, Moody’s analysts led by Vincent Gusdorf modeled a scenario in which valuations of AI-related companies fall 40% in the coming months. That figure should be read for what it is — an analytical scenario, not a base case — but it illuminates what the market did today.
“Tokenmaxxing”: the bill lands on the CFO’s desk
The other side of the day plays out inside companies, and it is the more instructive one for anyone managing a budget. As recently as this spring, “tokenmaxxing” denoted a virtue: maximizing token consumption, to the point of becoming a marker of internal modernity. Jensen Huang had set the yardstick: if a $500,000-a-year engineer isn’t burning $250,000 in tokens, something is wrong. Sam Altman said he was “excited to see what will happen with tokenmaxxing startups.” At Meta, an informal internal leaderboard christened “Claudeonomics” logged roughly 60 trillion tokens in a single month before being taken down.
Summer flipped the sign. According to the July 28 AP investigation, Uber consumed its entire 2026 AI budget in four months, then imposed a cap of $1,500 per month per coding tool. Salesforce is heading for an Anthropic bill of roughly $300 million for the year. Microsoft reined in individual Claude Code spending after some engineers ran up large personal bills. A consultant at Bain & Company sums up the arithmetic many organizations had not done: $200 per developer per month looks harmless, until you multiply it by a large group’s 20,000 developers — $48 million a year for a line item that was budgeted nowhere. And the trajectory is steeper than the level: for some large enterprises, the token bill is doubling almost every other month.
“It’s very easy to create something you don’t need with AI. As bills started to pile in, people realised that those new tools are quite expensive and you need to use them wisely.”
— Vincent Gusdorf, Head of AI Analytics, Moody’s Ratings
The rest of the cast says the same thing in different registers. Satya Nadella warned that tokenmaxxing is addictive and that the customer pays twice: in tokens, and through exposure of their proprietary data. Palantir’s Alex Karp judged that “something had gone completely wrong.” Raffi Krikorian, Mozilla’s chief technology officer, is blunt: “tokenmaxxing is a dumb thing.” The most operational remark, though, comes from the Bain consultant, and it fits in one line: not everything needs the most powerful model in the catalog.
The paradox: tokens have never been cheaper
Both ends have to be held at once, because this is where the costliest reasoning error hides. Unit prices have indeed fallen, massively. Between early 2025 and early 2026, API rates dropped roughly 80%. On July 24, Anthropic put Claude Opus 5 at the top of the Artificial Analysis index at $5 / $25 per million tokens, half the price of the Claude Fable 5 ($10 / $50) it dethroned. The cheapest paid model on the market, Ministral 3 3B, lists at $0.10 / $0.10 as of July 27. On paper, AI has never cost so little.
Volume did the rest, and it did far more. The shift toward agentic usage — a model that plans, calls tools, reviews its own output and starts again — consumes up to 1,000 times more tokens than a simple chatbot query. A price divided by five against a volume multiplied by a thousand produces a bill multiplied by two hundred. That is exactly what the finance departments quoted by AP describe: none of them got the sticker price wrong, all of them got the volume wrong. And this spending lands in a market that Gartner valued on May 19, 2026 at $2.59 trillion in worldwide AI spending for the year, up 47%, with more than 45% going to infrastructure alone.
The practical consequence is that the per-token price has ceased to be a management metric. The only one still operative is the cost per completed task, which accounts for the reasoning tokens actually burned. The gap is measurable: Artificial Analysis records $2.03 per task for Opus 5 at maximum effort against $2.75 for Fable 5, 26% less at equivalent intelligence — a gap that appears on no pricing page.
The July 27 answer isn’t more power, it’s routing
The clearest demonstration of this shift arrived the day before the crash, and it comes from Microsoft. On July 27, the company released MAI-Cyber-1-Flash, its first model dedicated to cybersecurity: a mixture-of-experts architecture with 137 billion parameters of which 5 billion are active, and a 256,000-token context window.
The point is not the model, it is the assembly. Inside the MDASH agentic system, this small model handles up to 90% of tasks, and calls GPT-5.4 only for the 10% that are genuinely hard. The announced result: 95.95% on the CyberGym benchmark, more than ten points ahead of competing configurations from Google and OpenAI, for 50% less than Microsoft’s previous combination (GPT-5.4, GPT-5.4 mini and GPT-5.3 Codex). In other words: better result, half the price, by changing not the model but the routing rule. Worth noting that Microsoft remains dependent on OpenAI for the hard segment — routing does not eliminate the frontier model, it stops paying for it on tasks that do not warrant it. In a single announcement, that is the refutation of tokenmaxxing.
Beneath the tokens: electricity, then compliance
Two costs still invisible on API invoices are surfacing. The first is physical. Inference now accounts for 80 to 90% of AI compute load at the major providers, and it is continuous. On the largest US power grid, PJM, the latest capacity auction came in 6.8 gigawatts short of the level needed to guarantee reliability at peak — the equivalent of nearly seven nuclear reactors — and supply costs rose more than 60%, with data centers accounting for roughly $6.3 billion of the $16.4 billion in those auctions (reading of July 15, 2026). Fortune noted on July 26 that these same data centers had until now tended to push electricity costs down, and that it is the scale of the investment program — put at some $7 trillion with no guaranteed demand behind it — that threatens to reverse the trend. Two figures circulating on this subject call for an explicit caveat: worldwide data center consumption in 2026, estimated at around 565 TWh (+26% year over year), and the count of 75 projects representing $130 billion postponed or canceled for lack of power, come from sector analyses we were unable to corroborate against a primary source.
The second cost is regulatory, and it has a date: this Sunday, August 2, 2026, Article 50 of the European AI Act enters into application. It imposes four transparency obligations: disclosing that a user is interacting with an AI, marking synthetic content in machine-readable form, informing people subject to biometric categorization, and disclosing deepfakes. Penalties can reach €15 million or 3% of worldwide turnover, whichever is higher. For a Swiss organization, the text applies not on grounds of territory but on grounds of market: as soon as content or systems reach users in the Union, the obligation follows. To this add the European Commission’s July 16 decision under the Digital Markets Act, which requires Google to open to competitors the Android system access reserved for Gemini — eleven features, available with Android 18 by August 1, 2027 at the latest — and to share its search data on fair terms from January 2027. Compliance is becoming a cost line on the same footing as compute.
Comparison table: API prices as of July 28, 2026
No pricing change has been published since the July 27 reading. The table below therefore serves as a unit-cost reference — bearing in mind that it is volume, not these figures, that made today’s news.
| Model | Input ($/M tokens) | Output ($/M tokens) | Note |
|---|---|---|---|
| Claude Fable 5 (Anthropic) | 10.00 | 50.00 | dethroned on the AA index July 24; data retention required |
| GPT-5.6 Sol (OpenAI) | 5.00 | 30.00 | AA index 59 |
| Claude Opus 5 (Anthropic) | 5.00 | 25.00 | #1 on the AA index (61); $2.03/task; Fast mode ×2.5 at twice the price |
| GPT-5.6 Terra (OpenAI) | 2.50 | 15.00 | launched July 9; some aggregators still show 1.25/7.50 |
| Kimi K3 (Moonshot, China, open-weight) | 3.00 | 15.00 | weights opened July 27 (1.4 TB); license unpublished; AA index 57 |
| Gemini 3.1 Pro (Google) | 2.00 | 12.00 | doubles beyond 200k context |
| Claude Sonnet 5 | 2.00 | 10.00 | launch price until August 31, then 3/15 |
| Gemini 3.6 Flash (Google) | 1.50 | 7.50 | launched July 21; −17% tokens/task |
| Grok 4.5 (xAI) | 2.00 | 6.00 | coding model; EU access still restricted |
| GPT-5.6 Luna (OpenAI) | 1.00 | 6.00 | launched July 9 |
| Inkling (Thinking Machines, open-weight) | 1.87 | 4.68 | Apache 2.0 license; multimodal |
| Meta Muse Spark 1.1 | 1.25 | 4.25 | Meta’s first paid API |
| LongCat-2.0 (Meituan, open-weight) | 0.75 | 2.95 | promo 0.30/1.20; MIT license |
| Gemini 3.5 Flash-Lite (Google) | 0.30 | 2.50 | launched July 21; very high throughput |
| Grok 4.3 / 4.20 (xAI) | 1.25 | 2.50 | 1M context |
| Mistral Large 3 | 0.50 | 1.50 | December 2025 rate (−75%); Batch API at −50% |
| DeepSeek V4 Pro | 0.435 | 0.87 | deepseek-reasoner alias removed July 24 |
| DeepSeek V4 Flash | 0.14 | 0.28 | 1M context; deepseek-chat alias removed July 24 |
| Ministral 3 3B (Mistral) | 0.10 | 0.10 | cheapest paid API recorded as of July 27 |
Mistral Large 3, Ministral 3 3B, DeepSeek V4, Opus 5, Fable 5 and Kimi K3 rates verified on July 28, 2026; the other lines are carried over from the July 27, 2026 reading, no pricing change having been published since. Caution: several third-party aggregators still publish stale rates — Mistral Large 3 is frequently shown at 2.00/6.00 when the official price has been 0.50/1.50 since December 2025. Promotional and cache rates can expire without notice.
What this means for a Swiss organization
Three conclusions, in the order in which they cost money.
First, the AI budget has to be a technical constraint, not an accounting line. Uber did not overshoot its budget because prices went up, but because nothing in its infrastructure prevented it from overshooting; the answer, a cap of $1,500 per tool per month, came after the fact. A cap per team, per project and per tool, enforced at the system level rather than on an expense report, is today the first mechanism to put in place — before even choosing a model.
Second, the saving is in the routing, not in the discount. MAI-Cyber-1-Flash puts numbers on it: 90% of tasks on a small model, 10% on the frontier model, and the bill halves at higher performance. Negotiating 10% off a rate will never make up for a task routed to a model ten times too expensive for it. That requires knowing, task by task, what each one actually consumes — which means measuring a cost per task, not a price per token.
Finally, what the Seoul session illustrates applies to vendors too. When a market assumption is revised, it is revised fast: −10.84% in one session, −29% over the month. An organization whose application foundation is welded to a single vendor absorbs those revisions — in pricing, in rate limits, in availability, and soon in Article 50 compliance — with no leverage. The right posture is neither to bring everything back in-house nor to entrust everything: it is to keep the measurement and the control, to know what each task costs and to be able to move it elsewhere without rewriting your foundation.
Control your AI costs instead of enduring them. Comoto OS orchestrates the market's best models, proprietary or open-weight, from a sovereign core hosted in Switzerland: every task routed to the model with the right cost/capability ratio, every token traced and measured as a real cost per task, caps enforced at the system level rather than discovered on the invoice, and the freedom to switch models without rewriting anything.
Discover Comoto OSBe the first to read our upcoming papers.
Some articles are reserved. Leave your email to unlock access to upcoming Meotis Research publications.




