DeepSeek dropped its flagship output token price to roughly $0.87 per million in May 2026. Claude Opus 4.6 still lists at $25. That 28x spread did not arrive overnight, and it did not arrive from a coordinated decision by five labs at once.
What triggered the 2026 Chinese AI token price war?
The 2026 Chinese AI token price war was triggered by DeepSeek making its 75% V4-Pro promotional discount permanent in late May 2026, not by a synchronized cut from five labs. DeepSeek launched V4-Pro on April 24, 2026 at $1.74 per million input and $3.48 per million output tokens, ran a 75% promotion, then around May 22-23 confirmed the discount would not roll back . List prices fell to about $0.435 per million input and $0.87 per million output, roughly one-quarter of launch .
Quick Answer: The 2026 price war began when DeepSeek made its V4-Pro 75% discount permanent in May 2026, pushing flagship output to ~$0.87/M tokens. Other labs followed over days to weeks, not a coordinated five-lab move. Against Claude Opus 4.6 at $25/M output, that is a 28x spread.
This is the second wave, not the first. DeepSeek's V2 paper, submitted May 7, 2024, described a 236B-parameter Mixture-of-Experts model with 21B active parameters and a 93.3% smaller KV cache. Those efficiency gains underwrote unusually low inference prices and helped ignite the original 2024 cascade . Both episodes share a pattern: DeepSeek triggers, others react.
The popular "five labs slashed prices simultaneously by up to 99%" framing is inaccurate. Cuts clustered over days to weeks behind the DeepSeek trigger rather than landing as one pact. The 99% figure came specifically from Xiaomi, which on May 27, 2026 cut prices across its MiMo-V2.5 series by up to 99% and gave existing customers five to eight times more usable credits . That "up to 99%" is a ceiling for one aggressive mover and its lightweight tiers, not an average across every lab and model. The rest of this analysis breaks down who actually cut, who held, and what each lab charges now.
Why DeepSeek made the 75% V4-Pro cut permanent

DeepSeek made its discount permanent because the low price was always architecturally affordable, not a temporary loss-leader. The company launched its flagship V4-Pro on April 24, 2026 at $1.74 per million input tokens and $3.48 per million output tokens, with a 75% promotional discount attached from day one . On May 22-23, 2026 it confirmed the discount would not roll back, settling list prices at roughly $0.435 input and $0.87 output per million tokens, about one-quarter of launch pricing .
Quick Answer: DeepSeek made its 75% V4-Pro discount permanent on May 22-23, 2026, fixing list prices at ~$0.435 input / $0.87 output per million tokens. The model's Mixture-of-Experts design (21B active of 236B parameters, a 93.3% smaller KV cache, and 5.76x higher throughput) makes that price structurally sustainable, not a temporary promotion.
The most consequential detail for builders is the cache-hit rate. DeepSeek quotes input on a KV-cache hit near $0.003625 per million tokens, two orders of magnitude below the cache-miss rate . That directly rewards cache-stable prompt patterns: agents with fixed system prompts, RAG loops that reuse retrieved context, and repeated-structure completions all hit the cache far more often than ad-hoc one-off calls. Prompt architecture, not just model choice, now moves your bill.
"The fixed prefix is the cheapest token you'll ever send. Design your agent so the system prompt and tool definitions never shift between turns," argues the context-engineering walkthrough on prompt-stable agent design (video: Robert Ta).
The pricing has architectural grounding. DeepSeek's MoE lineage activates only 21B of 236B total parameters per token, cuts the KV cache by 93.3% versus the prior dense generation, and reaches 5.76x higher maximum generation throughput . Fewer active parameters and a smaller cache mean lower memory and compute per served token. Those are the inference economics that let a sub-$1 output price hold without a permanent subsidy.
Coverage framed the move as a strategic pivot "from building the smartest models to delivering the lowest-cost services" . For developers, the practical takeaway is concrete: a stable, cache-friendly prompt layout against V4-Pro can land effective input costs near the cache-hit floor, turning multi-step agent loops that were prohibitively expensive on US flagships into routine workloads.
The cascade after Xiaomi: who cut and who didn't
Xiaomi turned DeepSeek's permanent cut into a sector event. From May 27, 2026, it reduced API prices across the entire MiMo-V2.5 series by up to 99% and handed existing customers five to eight times more usable credits for the same spend . That "up to 99%" is a ceiling on the most aggressive tiers, not an average. It set the framing for everything that followed.
The frontier it defined is the sub-$1-per-million-output line, and labs split cleanly on whether they crossed it. The movers:
- Xiaomi MiMo V2 Flash: roughly $0.09 input / $0.29 output per million tokens, accounting for about 21.1% of OpenRouter traffic (~4.21 trillion weekly tokens) .
- Alibaba Qwen 3.5 Flash: quoted near $0.065 input / $0.26 output per million, undercutting even Xiaomi on input .
- Baidu Ernie: cut its latest lineup and pushed some lightweight variants to free .
Two notable holdouts stayed above the line. Zhipu's GLM 5.1 sits near $4.40 per million output and Moonshot's Kimi K2.6 near $4.00, both visibly above DeepSeek's roughly $0.87 output and Xiaomi's $0.29 . For a developer routing high-volume traffic, that gap is a routing decision, not a rounding error: a multi-step agent loop that costs cents on Qwen Flash can cost dollars on GLM 5.1 at the same token count.
The lab-specific pattern matters because the cascade is not a sector-wide race to zero. Earlier in 2026 Zhipu raised API prices by 83%, and on March 18, 2026 Alibaba raised AI compute prices by up to 34% when demand outran its infrastructure . Those reversals are the tell: where the cut was promotional or capacity-constrained, prices snapped back the moment utilization tightened. The takeaway for builders: treat each provider's pricing as an independent, capacity-bound signal. Confirm live rates per model rather than assuming a uniform downward trend across "the five."
What each Chinese AI lab charges in 2026: token price table

In mid-2026, China's flagship and lightweight LLM tiers cluster well below comparable US list prices, but the spread between labs is wide, from DeepSeek's sub-dollar flagship output to Zhipu and Moonshot rates that sit closer to US territory. The table below consolidates the published per-million-token rates from the research, alongside the two US reference points developers most often benchmark against. Treat it as a snapshot for sizing a workload, not a contract: several of these figures move on promotional timelines.
| Lab | Flagship / variant | Input $/M | Output $/M | Context | Cache-hit input $/M |
|---|---|---|---|---|---|
| DeepSeek | V4-Pro | $0.435 | $0.87 | — | $0.003625 |
| DeepSeek | V4-Flash | $0.14 | $0.28 | 1M | — |
| Xiaomi | MiMo V2 Flash | $0.09 | $0.29 | — | — |
| Alibaba | Qwen 3.5 Flash | $0.065 | $0.26 | — | — |
| Baidu | ERNIE-Lite-Pro-128K | RMB 0.0002 / 1K | RMB 0.0004 / 1K | 128K | — |
| Zhipu | GLM 5.1 | — | ~$4.40 | — | — |
| Moonshot | Kimi K2.6 | — | ~$4.00 | — | — |
| OpenAI (ref.) | GPT-5.4 | $2.50 | $15.00 | — | — |
| Anthropic (ref.) | Claude Opus 4.6 | $5.00 | $25.00 | — | — |
DeepSeek V4-Pro anchors the aggressive end: $0.435 input and $0.87 output per million tokens after the permanent discount, with a cache-hit input rate quoted near $0.003625 per million, a roughly 120x gap between cache-miss and cache-hit input that directly rewards prompt-stable agent designs . The cheaper V4-Flash lists at $0.14 / $0.28 with a 1M-token context window .
The lightweight tiers go lower still. Xiaomi's MiMo V2 Flash sits near $0.09 / $0.29 and Alibaba's Qwen 3.5 Flash near $0.065 / $0.26 per million tokens . Baidu publishes ERNIE-Lite-Pro-128K in RMB-per-1K units (RMB 0.0002 input and RMB 0.0004 output per 1K tokens) and made some lightweight versions free outright .
Not every Chinese flagship undercuts. Zhipu's GLM 5.1 (~$4.40 output) and Moonshot's Kimi K2.6 (~$4.00) look uncompetitive against DeepSeek's ~$0.87, though both still land below GPT-5.4 at $15.00 and Claude Opus 4.6 at $25.00 . One operational caveat before you commit a workload: these are list figures gathered across reports, and unit conventions differ (per-million vs. per-1K, cache-hit vs. cache-miss). Confirm every rate against each lab's official pricing page, since live catalogs have already overwritten several 2026 before/after tables .
Chinese AI at 45% of OpenRouter: how the supply equation tipped
The price war's clearest demand-side signal is OpenRouter's traffic mix: Chinese models held under 2% of token consumption in early 2025 and exceeded 45% by April 2026 . That is the fastest supply-side concentration shift in the platform's history, and it crossed the halfway line fast. Chinese models reportedly overtook US rivals in aggregate OpenRouter consumption around February 2026 . Cheaper tokens did not just win benchmarks; they captured routing decisions at scale.
One provider concentrated much of that gain. As of Q2 2026, Xiaomi alone accounted for roughly 21.1% of OpenRouter traffic, about 4.21 trillion weekly tokens and roughly three times OpenAI's share on the same platform . The field also consolidated to roughly ten serious providers after the cascade: Xiaomi, Alibaba, Zhipu, DeepSeek, Moonshot, MiniMax, StepFun, ByteDance, Baidu and Tencent .
| OpenRouter metric (Q2 2026) | Figure |
|---|---|
| Chinese-model share, early 2025 | Under 2% |
| Chinese-model share, April 2026 | Over 45% |
| Xiaomi share of platform traffic | ~21.1% (~4.21T weekly tokens) |
| Xiaomi vs. OpenAI share | ~3x larger |
| Serious providers post-consolidation | ~10 |
The demand is not evenly distributed across use cases. It clusters in high-volume, price-elastic application categories where the per-token gap compounds at scale: chatbots, customer-service automation, coding tools, and retrieval-augmented generation (RAG) pipelines feeding multi-step agents . These are exactly the workloads that emit millions of tokens per session, so a 4-8x output-price gap versus US flagships turns directly into a routing default rather than a rounding error. The economic logic also rewards prompt-stable, cache-friendly architectures. This is the same context-engineering discipline that lets agent loops reuse cached prefixes instead of re-paying for them on every step (video: Robert Ta).
"DeepSeek started a price war. Now every Chinese AI lab wants in, shifting the contest from building the smartest models to delivering the lowest-cost services," — industry analysis (source: Technology.org, 2026-02).
The caveat for anyone reading these numbers as a verdict: OpenRouter share measures token volume routed through one aggregator, not enterprise revenue or production parity. It tracks where price-sensitive, high-throughput traffic flows, a leading indicator of commoditization, not proof that a cheaper token is an equivalent one.
The sustainability question: subsidized expansion or viable margin?

The sustainability question splits into two parts: are near-zero tiers profitable operations or developer-acquisition subsidies, and what reprices when the subsidy ends? The honest answer is that both forces are running at once. Baidu cut its latest Ernie and made some lightweight versions free , and Xiaomi handed existing MiMo-V2.5 customers five to eight times more usable credits at the same price from May 27, 2026, moves that read as customer capture, not margin .
The counter-signal is architectural, and it matters because it sets a floor that survives any promotion ending. DeepSeek's Mixture-of-Experts design activates only 21B of 236B parameters per token, cuts the KV cache by 93.3%, and reports 5.76x higher maximum generation throughput versus its dense predecessor . Those are structural unit-economics gains, not a discount with an expiry date, which is why DeepSeek could make the V4-Pro 75% cut permanent rather than letting it roll back . A lab with a genuinely lower cost-to-serve can hold a low price; a lab subsidizing one cannot.
And not every lab held. Alibaba raised AI compute prices by up to 34% on March 18, 2026 when demand outran infrastructure, and Zhipu raised API prices 83% earlier in the year . Those reversals are the tell: where the cut was promotional or capacity-constrained, prices snapped back the moment utilization tightened. As one context-engineering practitioner frames the builder's exposure, "the cheapest token only stays cheap while the provider wants your traffic; your architecture has to assume the rate card moves" (video: Robert Ta).
The likely equilibrium is bifurcation rather than collapse. Near-free tiers persist as onboarding tools, while high-concurrency, SLA-backed capacity reprices upward toward real cost. For developers, the early-warning instrument is not the headline output rate but the published rate limits at the free tier. Watch for tightening request-per-minute caps, lower context ceilings, or new throughput throttles at the free level; those changes usually precede a price move and signal that a given lab is shifting a model from acquisition to monetization. Pair that with DeepSeek's cache-hit input rate near $0.003625 per million tokens , which rewards prompt-stable agent designs, and the defensible read is that structural-cost leaders sustain low pricing while subsidy-driven tiers are the ones to stress-test before you build on them.
The quality parity problem: when cheaper tokens aren't equivalent
A cheaper token is not an equivalent token. Per-token price captures one dimension of cost, not total fit: context window length, tool-call support, concurrency caps, safety-filtering behavior, and data-residency terms vary enough across Chinese labs that two endpoints at the same headline rate can deliver very different production economics. DeepSeek-V4-Flash, for example, advertises a 1M-token context length alongside its $0.14 input / $0.28 output per million pricing , while lightweight tiers from Baidu's Qianfan such as ERNIE-Lite-Pro-128K trade context and capability for a lower sticker . The token rate tells you almost nothing about which of those fits your workload.
Benchmark verification is the second gap. Independent leaderboard coverage is sparse for several mid-tier Chinese labs, and the "five labs, simultaneous, up to 99%" framing already shows how vendor narratives compress messy reality. The 99% figure traces to Xiaomi's MiMo-V2.5 cuts from May 27, 2026 and applied to specific aggressive movers and lightweight tiers, not every model . Treat a vendor's quoted benchmark the same way: cross-check it against third-party evals before committing a high-stakes workload, rather than repeating the claim.
Operational terms are the third. Latency profiles and uptime SLAs are not consistently published across these providers, a real problem for latency-sensitive or regulated applications, where Chinese cloud geography adds round-trip overhead that no price sheet reflects. The same caution applies to cache economics: DeepSeek's cache-hit input rate near $0.003625 per million tokens only materializes if your prompt structure is actually cache-friendly.
The practical takeaway: run your real workload against candidate endpoints instead of trusting generic benchmarks, and confirm cache pricing applies to your specific prompts before projecting any savings. Context-engineering practitioners make the same point: stable, cache-aligned prompts are what convert headline rates into actual margin (video: Robert Ta). When Chinese flagships undercut GPT-5.4 ($2.50/$15.00) and Claude Opus 4.6 ($5.00/$25.00) by 4-8x on output , the savings are large enough to justify a week of measurement before you migrate. They are not large enough to skip it.
Frequently asked questions
Did DeepSeek lock in the 75% V4-Pro discount long-term?
Yes. On May 22-23, 2026, DeepSeek confirmed the V4-Pro promotional discount would not roll back, making it permanent . The standing list price is roughly $0.435 per million input tokens and $0.87 per million output tokens, down from a $1.74 / $3.48 launch on April 24, 2026 . Cache-hit input is quoted near $0.003625 per million tokens . Confirm current rates on DeepSeek's official pricing page, since list prices update without notice.
How do Chinese AI token prices compare to US labs in 2026?
Chinese flagships undercut their US equivalents by roughly 2.5-5x on input and 4-8x on output as of mid-2026 . DeepSeek V4-Pro outputs cost $0.87 per million tokens against GPT-5.4 at $15.00 and Claude Opus 4.6 at $25.00 . The gap is not uniform, though: Zhipu's GLM 5.1 (~$4.40 per million output) and Moonshot's Kimi K2.6 (~$4.00) sit well above DeepSeek and above some cheaper US tiers .
Does Xiaomi's 99% slash apply across all its variants?
No. The cut applies to the MiMo-V2.5 series from May 27, 2026, and the "up to 99%" figure is a ceiling for the most aggressive reductions, not a uniform average across every variant . Existing customers also received five to eight times more usable credits for the same price . In concrete terms, the MiMo V2 Flash tier landed near $0.09 input and $0.29 output per million tokens .
Are Chinese AI token prices sustainable or a temporary subsidy?
The signals are mixed. DeepSeek's Mixture-of-Experts design (236B total parameters with only 21B active per token, a 93.3% smaller KV cache, and 5.76x higher generation throughput in the V2 paper) provides genuine structural cost advantages that underwrite low inference prices . But near-zero tiers from some labs look more like developer-acquisition subsidies than viable margins. Alibaba reportedly raised AI compute prices by up to 34% on March 18, 2026, and Zhipu lifted API prices 83% earlier in the year when economics demanded it . The "race to zero" framing is oversimplified.
What should you evaluate beyond token price when choosing a Chinese AI endpoint?
A cheaper token is not an equivalent token. Before migrating, weigh context window length, tool-call support, concurrency and rate limits, safety-filter behavior, data residency, and uptime SLAs . Verify benchmark claims against independent leaderboards rather than vendor numbers, since generic rankings for mid-tier Chinese labs are often thin or unverified by third parties . The reliable test is to run your own target workload (chatbot, RAG, or multi-step agent) and measure cost-per-resolved-task, not headline price.