DeepSeek’s V4-Flash costs three cents per benchmark test on Artificial Analysis’s Intelligence Index, more than 100 times cheaper than Anthropic’s Claude Fable 5, which the San Francisco research firm clocks at $3.15. The official rack rate is $0.14 per million input tokens and $0.28 per million output tokens. On the same benchmark suite, Moonshot AI’s Kimi K3 costs 86 cents per test and OpenAI’s GPT-5.6 Sol runs $1.86.
The intelligence gap is narrower than the price gap suggests. V4-Flash scored 50 out of 100 on the nine-benchmark index that averages coding, reasoning, and workplace tasks, sitting one point behind Meta’s Muse Spark 1.1 and Z.AI’s GLM-5.2, and seven behind Kimi K3 at 57. Claude Opus 5, Claude Fable 5, and GPT-5.6 all sit nine or more points higher. Axios puts the practical spread another way: roughly a 99% discount on comparable output versus Claude Opus 4.8.
That framing is what forced OpenAI’s hand. On July 30, the company cut GPT-5.6 Luna pricing by 80%, to $0.20 per million input tokens and $1.20 per million output. GPT-5.6 Terra came down 20%, to $2 and $12. GPT-5.6 Sol Standard held firm at $5 and $30, with a new Sol Fast mode priced at double the Standard rate in exchange for a claimed 2.5× throughput. Anthropic hasn’t blinked: Claude Opus 5 ships at the same $30 combined rate as Opus 4.8.
The Chinese pricing wave isn’t an isolated stunt. VentureBeat reports that Chinese models now capture 46% of U.S. enterprise token usage on OpenRouter, a share that would’ve read as implausible eighteen months ago and now reads as a structural fact developers route around cost, and the cheapest competent model wins the API call.
There’s a familiar shape to this. The 2015–2017 collapse in cloud-storage pricing across AWS, Azure, and Google Cloud followed the same pattern: a challenger with a different cost structure prices aggressively, the incumbent holds its premium tier while quietly slashing the commodity tier, and within two years the entire category resets around the lower anchor. Storage became a loss leader for compute. Inference is on track to become a loss leader for something else, agents, tooling, model routing, whatever the next margin layer turns out to be.
What’s striking is how quickly the elite narrative has shifted from capability to cost. A year ago the interesting question was whether frontier labs could keep scaling. Now it’s whether anyone can charge for the results. Sam Altman, per SCMP’s read on the Luna cut, is the one who blinked. The Intelligence Index still favors Anthropic and OpenAI at the top. The purchase orders increasingly don’t.
Sources
- https://www.reuters.com/technology/artificial-intelligence/deepseeks-new-ai-model-is-by-far-cheapest-well-known-models-run-research-firm-says-2026-08-03/
- https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html
- https://www.axios.com/2026/08/01/deepseek-model-cheap-ai-price-war
- https://www.scmp.com/tech/tech-trends/article/3362568/openai-blinks-face-off-with-chinese-rivals-drops-pricing-some-models-80
- https://venturebeat.com/technology/ai-price-wars-openai-cuts-gpt-5-6-luna-prices-by-80-as-model-competition-shifts-toward-cost