The Token Price War: From “1 Yuan per Million Tokens” to K-Shaped Differentiation
In May 2024, DeepSeek-V2 reset the market with a shocking price: 1 RMB per million input tokens. Within a year, Alibaba’s Tongyi cut prices by 97%, ByteDance’s Doubao reached 0.0008 RMB per thousand tokens, and “1 yuan buys 1 million tokens” became the consumer tagline. Then in 2026, the market split: Zhipu raised prices three times (cumulative ~83%) and watched its OpenRouter ranking collapse from #3 to #17, while DeepSeek cut prices to 25% of original and pushed cache-hit pricing to 0.025 RMB per million tokens. This is the story of how China’s token pricing evolved — and what every foreign procurement team must understand before signing a contract.
📋 What’s Inside
1. Phase 1: The Blanket Price War (2024–2025)
The price war began in May 2024 when DeepSeek-V2 launched with API pricing of 1 RMB per million input tokens and 2 RMB per million output tokens — roughly 1% of OpenAI’s GPT-4 Turbo pricing at the time. This single move detonated a 12-month price carnage that reshaped the entire Chinese AI services market.
The cascade was immediate and brutal:
- May 2024: DeepSeek-V2 sets the 1 RMB/million token benchmark — 1/70 to 1/35 of GPT-4 Turbo’s input pricing.
- May 2024 (same month): Baidu immediately makes ERNIE Speed and ERNIE Lite completely free; Alibaba cuts Tongyi Qwen-Long input to 0.0005 RMB per thousand tokens (97% reduction, 1/400 of GPT-4).
- Mid-2024: ByteDance launches Doubao at 0.0008 RMB per thousand tokens — advertised as “99.3% cheaper than industry.”
- January 2025: Alibaba cuts Qwen-VL vision model pricing by over 80%; DeepSeek-R1 launches at 1 RMB/million input (cache hit), far below OpenAI o1’s 438 RMB/million output.
- By mid-2025: Average API prices across major Chinese models have fallen more than 90% year-over-year.
2. Phase 2: K-Shaped Differentiation (2026–)
By early 2026, the market bifurcated into a “K-shape”: one branch of vendors raised prices and lost share; the other cut prices and won. The two flagship cases define the era:
📈 Zhipu: The Failed Price Hike
Mar 16: GLM-5-Turbo API +20%
Apr 8: GLM-5.1 API +10%
OpenRouter Rank: #3 (mid-Feb) → #17 (mid-Apr)
“Regressing token prices to normal commercial value” — Zhipu CEO Peng Zhang
📉 DeepSeek: The Permanent Cut
Cache-miss input: 3 RMB/million tokens
Output: 6 RMB/million tokens
Global comparison: ~1/60 of Fable 5, ~1/30 of Claude Opus 5
Result: #1 Chinese model by call volume
Zhipu’s experience is the cautionary tale of the year. When GLM-5 launched in mid-February 2026, it ranked #3 globally on OpenRouter — the highest-ranked Chinese model on the platform. After three rounds of cumulative 83% price increases, by mid-April its ranking had collapsed to #17. Developers voted with their wallets. As Zhipu CEO Peng Zhang stated at the Zhongguancun Forum: “The bottleneck is compute, not customers” — yet the market disagreed.
Meanwhile, DeepSeek’s strategy proved the opposite thesis. By permanently cutting V4-Pro to 25% of its original price (announced May 2026, effective after the 2.5-discount promotional period ended), DeepSeek locked in developer mindshare and sustained its position as China’s most-called model. The cache-hit input price of 0.025 RMB per million tokens — approximately 34× cheaper than GPT-5.5’s equivalent — created a structural cost advantage that competitors could not match.
3. The Cache Economy: Approaching Zero Marginal Cost
The most important pricing innovation in China’s token market is not the headline output price — it is the cache-hit differential. When a user sends a prompt that has been seen before, the system reuses pre-computed results instead of re-running inference. This makes the marginal cost of a cache-hit token approach zero.
| Model | Input (Cache Hit) | Input (Cache Miss) | Output | Cache Discount |
|---|---|---|---|---|
| DeepSeek V4-Pro | 0.025 RMB | 3 RMB | 6 RMB | 120× cheaper |
| DeepSeek V4-Flash | 0.02 RMB | 1 RMB | 2 RMB | 50× cheaper |
| GPT-5.5 (International) | ~210 RMB output / million tokens | ~1/34 of DeepSeek V4-Pro output | ||
This creates a fundamental economic truth: for applications with high cache hit rates (chatbots with system prompts, RAG systems, repetitive Agent tasks), the effective token cost is approaching zero. A system with 90% cache hits paying 0.025 RMB for those tokens and 3 RMB for the remaining 10% pays an effective blended rate of roughly 0.32 RMB per million input tokens — cheaper than the price of the electricity consumed to generate them.
4. Time-of-Use Pricing: Tokens Like Electricity
In a landmark move, DeepSeek announced in 2026 that it would introduce time-of-use (TOU) pricing — the first major AI vendor globally to apply electricity-market logic to token pricing. Under the proposed scheme:
| Time Period (Beijing Time) | V4-Pro Input (Cache Hit) | V4-Pro Input (Cache Miss) | V4-Pro Output |
|---|---|---|---|
| Off-Peak (all other hours) | 0.025 RMB | 3 RMB | 6 RMB |
| Peak: 9:00–12:00 & 14:00–18:00 | 0.05 RMB (2×) | 6 RMB (2×) | 12 RMB (2×) |
This innovation reflects a deeper reality: token production is a spatio-temporal scarce resource. During peak hours, GPU clusters in western China run at maximum utilization; during off-peak hours, vast amounts of compute sit idle. By doubling prices during peak windows, DeepSeek converts idle capacity into incremental revenue while incentivizing customers to shift flexible workloads to off-peak hours — exactly as electricity markets have done for a century.
5. The International Price Gap & Export Arbitrage
The most strategically significant pricing fact for foreign firms is the persistent 10× to 34× gap between Chinese and international model output prices. The top five Chinese models’ output pricing stands at 1/10 to 1/34 of global giants like GPT-5.5.
This gap is not closing — it is widening. While Chinese vendors like Zhipu attempt to raise prices (and face ranking collapses), the leading players (DeepSeek, Doubao) are aggressively cutting prices to capture global developer mindshare. The economic logic is inexorable:
- Cost advantage: Cheaper green power in western China, domestic chip economics (Huawei Ascend at 41% domestic share), and inference engine optimization combine to push Chinese token costs structurally below global equivalents.
- Strategic intent: Chinese vendors are explicitly using low pricing as a tool for overseas expansion. The 3–5× international price arbitrage means a Chinese AI vendor can profitably sell tokens to European or Southeast Asian customers at prices that undercut local providers by 70–90%.
- Risk for foreign partners: A vendor whose entire business model depends on unsustainable domestic pricing may not survive the next 12–18 months as the market consolidates. Conversely, a vendor with healthy overseas revenue (3–5× margins) has structural resilience.
6. Contract Negotiation in the Token Era
For overseas procurement teams, the token pricing landscape of 2026 demands a fundamentally different negotiation approach than even 12 months ago. Here is a practical framework:
6.1 What to Watch in Every Quote
- Cache-hit rate commitments: Request the vendor’s cache-hit statistics for workloads similar to yours. A vendor refusing to disclose this metric is hiding their true effective pricing.
- Tiered pricing schedules: Verify whether quoted rates cover input (cache hit vs. miss separately), output, and any premium tiers. A single “per million token” number is meaningless without this breakdown.
- Peak/off-peak differentials: Confirm which hours are classified as peak in the vendor’s TOU schedule and whether your workloads can be scheduled to avoid them. Get off-peak pricing locked in contractually.
- Price escalation clauses: Given that Morgan Stanley data shows 48–80% price increases in just 12 months, contracts without price protection expose you to unilateral hikes. Negotiate annual caps or formula-based adjustments.
- Currency and payment terms: RMB pricing exposes foreign buyers to FX risk. Explore whether the vendor offers USD/EUR denominated contracts with fixed exchange rate windows.
6.2 Your Negotiation Leverage
The market remains firmly a buyer’s market for foreign firms bringing meaningful volume. Vendors are still fighting for call volume to achieve scale economies. This means:
- Multi-year committed use discounts: Lock in 12–24 month pricing with volume commitments. Vendors will trade price certainty for demand certainty.
- Custom SLAs: Demand uptime guarantees, latency targets, and priority queue access — these are negotiable when you commit to volume.
- Multi-vendor diversification: Use the threat of switching (DeepSeek ↔ Alibaba ↔ Huawei Cloud) as ongoing leverage. Vendors know lock-in is fragile in this market.
- Capacity reservations: For mission-critical workloads, reserve dedicated capacity at fixed rates — this protects you from both price hikes and capacity shortages during peak demand.
6.3 Red Flags in Token Pricing
Vendor quotes a single blended rate without separating cache-hit/cache-miss input and output. This obscures the true effective pricing and prevents you from optimizing your workload.
Contract allows unilateral price increases with minimal notice. In a market where prices swung 80%+ in 12 months, unprotected contracts are financially dangerous.
Vendor quotes prices below their own cost structure (verifiable through official enterprise credit reports). This is a cash-burn strategy that signals imminent price hikes or business failure.
Vendor refuses to specify TOU schedules or won’t guarantee off-peak access. As DeepSeek’s model proves, peak pricing is the future — ambiguity here means budget uncertainty.
🎯 Your 3-Step Action Plan for Token Pricing Negotiation
Verify before you sign. Commission an independent professional enterprise credit report on any Chinese AI vendor you are considering. Confirm their financial sustainability, revenue composition, and actual pricing model — because in a market transforming 80%+ per year, trust must be built on evidence.
Negotiate the effective rate, not the headline. Demand cache-hit statistics for your workload class, lock in off-peak pricing, secure multi-year price caps, and get peak-hour definitions in writing. The blended effective rate — not the marketing number — is what you pay.
Diversify and protect. Never depend on a single vendor. Design your China AI architecture for multi-vendor interoperability, reserve dedicated capacity for mission-critical workloads, and build contractual exit clauses. The K-shaped market will claim casualties — make sure you are not one of them.
At ChinaBizInsight, we help overseas firms navigate exactly these complexities. Our professional enterprise credit reports give you verified intelligence on Chinese AI companies’ pricing models, financial health, and business sustainability — because in a market where prices swing 80%+ annually, evidence is your only protection.
📚 References & Data Sources
DeepSeek official announcements (May 2026) — V4-Pro API permanently reduced to 25% of original pricing: cache-hit input 0.025 RMB, cache-miss input 3 RMB, output 6 RMB per million tokens; TOU pricing preview with 2× peak multipliers during 9:00–12:00 and 14:00–18:00 Beijing Time [3,9,12,18](@ref).
Tencent News (April 13, 2026) and Sina Finance (May 7, 2026) — Zhipu GLM-5 three price increases cumulative ~83%; OpenRouter ranking collapse from #3 (mid-Feb) to #17 (mid-Apr); Zhipu CEO Peng Zhang’s statement on “regressing token prices to normal commercial value” [5,14](@ref).
IMA Knowledge / industry analysis (July 2026) — DeepSeek-V2 May 2024 pricing at 1 RMB/1M input, 2 RMB/1M output (~1% of GPT-4 Turbo); Alibaba Tongyi Qwen-Long 97% price cut to 0.0005 RMB/thousand tokens; Doubao at 0.0008 RMB/thousand tokens (“99.3% cheaper than industry”) [1](@ref).
Morgan Stanley China AI Pricing Tracker (Q2 2026, reported August 2026) — Average Chinese API input price 4.9 RMB/million tokens (+48% vs. Q1 2025), output 21.9 RMB/million tokens (+80% vs. Q1 2025) [13](@ref).
Alibaba Cloud Developer Community (August 2026) — DeepSeek V4-Pro formal release: output at 6 RMB/million tokens = 1/60 of Fable 5 (~360 RMB), 1/30 of Claude Opus 5 (~180 RMB) [15](@ref).
Zhongshang Industrial Research Institute, China Token Factory Development White Paper 2026 — Chinese top-5 model output pricing at 1/10 to 1/34 of global giants; 3–5× arbitrage when selling tokens to overseas customers; cache-hit pricing approaching “milli-yuan” level with marginal cost near zero.
Note: All statistics current as of August 2026. DeepSeek’s announced TOU pricing is a “planned” mechanism as of August 2026; V4-Pro current formal pricing remains at the permanently reduced 25% level. Forward pricing is subject to vendor announcements and market conditions.
ChinaBizInsight
Your strategic bridge to transparent business in China.