ChinaBizInsight

CHINA AI INFRASTRUCTURE SERIES · PART 5

The Token Price War: From “1 Yuan per Million Tokens” to K-Shaped Differentiation

In May 2024, DeepSeek-V2 reset the market with a shocking price: 1 RMB per million input tokens. Within a year, Alibaba’s Tongyi cut prices by 97%, ByteDance’s Doubao reached 0.0008 RMB per thousand tokens, and “1 yuan buys 1 million tokens” became the consumer tagline. Then in 2026, the market split: Zhipu raised prices three times (cumulative ~83%) and watched its OpenRouter ranking collapse from #3 to #17, while DeepSeek cut prices to 25% of original and pushed cache-hit pricing to 0.025 RMB per million tokens. This is the story of how China’s token pricing evolved — and what every foreign procurement team must understand before signing a contract.

1. Phase 1: The Blanket Price War (2024–2025)

The price war began in May 2024 when DeepSeek-V2 launched with API pricing of 1 RMB per million input tokens and 2 RMB per million output tokens — roughly 1% of OpenAI’s GPT-4 Turbo pricing at the time. This single move detonated a 12-month price carnage that reshaped the entire Chinese AI services market.

1 RMB
DeepSeek-V2 input price per million tokens (May 2024)
-97%
Alibaba Tongyi Qwen-Long input price cut
0.0008 RMB
Doubao Pro-32K per thousand tokens
-90%+
Average API price decline for major models (2025)

The cascade was immediate and brutal:

  • May 2024: DeepSeek-V2 sets the 1 RMB/million token benchmark — 1/70 to 1/35 of GPT-4 Turbo’s input pricing.
  • May 2024 (same month): Baidu immediately makes ERNIE Speed and ERNIE Lite completely free; Alibaba cuts Tongyi Qwen-Long input to 0.0005 RMB per thousand tokens (97% reduction, 1/400 of GPT-4).
  • Mid-2024: ByteDance launches Doubao at 0.0008 RMB per thousand tokens — advertised as “99.3% cheaper than industry.”
  • January 2025: Alibaba cuts Qwen-VL vision model pricing by over 80%; DeepSeek-R1 launches at 1 RMB/million input (cache hit), far below OpenAI o1’s 438 RMB/million output.
  • By mid-2025: Average API prices across major Chinese models have fallen more than 90% year-over-year.
The “1 Yuan per Million Tokens” era: The consumer-facing slogan captured the moment perfectly. For the first time in computing history, intelligent output had become cheaper than the electricity required to generate it. But this was never sustainable — and everyone knew it.

2. Phase 2: K-Shaped Differentiation (2026–)

By early 2026, the market bifurcated into a “K-shape”: one branch of vendors raised prices and lost share; the other cut prices and won. The two flagship cases define the era:

📈 Zhipu: The Failed Price Hike

Raised prices 3 times, cumulative ~83%
Feb 12: GLM-5 launch + Coding Plan restructure (30%+ increase)

Mar 16: GLM-5-Turbo API +20%

Apr 8: GLM-5.1 API +10%

OpenRouter Rank: #3 (mid-Feb) → #17 (mid-Apr)

“Regressing token prices to normal commercial value” — Zhipu CEO Peng Zhang

📉 DeepSeek: The Permanent Cut

Cut V4-Pro prices to 25% of original
Cache-hit input: 0.025 RMB/million tokens

Cache-miss input: 3 RMB/million tokens

Output: 6 RMB/million tokens

Global comparison: ~1/60 of Fable 5, ~1/30 of Claude Opus 5

Result: #1 Chinese model by call volume

Zhipu’s experience is the cautionary tale of the year. When GLM-5 launched in mid-February 2026, it ranked #3 globally on OpenRouter — the highest-ranked Chinese model on the platform. After three rounds of cumulative 83% price increases, by mid-April its ranking had collapsed to #17. Developers voted with their wallets. As Zhipu CEO Peng Zhang stated at the Zhongguancun Forum: “The bottleneck is compute, not customers” — yet the market disagreed.

Meanwhile, DeepSeek’s strategy proved the opposite thesis. By permanently cutting V4-Pro to 25% of its original price (announced May 2026, effective after the 2.5-discount promotional period ended), DeepSeek locked in developer mindshare and sustained its position as China’s most-called model. The cache-hit input price of 0.025 RMB per million tokens — approximately 34× cheaper than GPT-5.5’s equivalent — created a structural cost advantage that competitors could not match.

The Morgan Stanley data point that explains everything: According to Morgan Stanley’s Q2 2026 tracking of major Chinese vendors’ official pricing, the average API input price in China rose to 4.9 RMB per million tokens, and output to 21.9 RMB per million tokens — representing increases of approximately 48% and 80% respectively compared to Q1 2025 (3.3 RMB input, 12.2 RMB output). The K-shape is not theoretical — it is already in the data.

3. The Cache Economy: Approaching Zero Marginal Cost

The most important pricing innovation in China’s token market is not the headline output price — it is the cache-hit differential. When a user sends a prompt that has been seen before, the system reuses pre-computed results instead of re-running inference. This makes the marginal cost of a cache-hit token approach zero.

Model Input (Cache Hit) Input (Cache Miss) Output Cache Discount
DeepSeek V4-Pro 0.025 RMB 3 RMB 6 RMB 120× cheaper
DeepSeek V4-Flash 0.02 RMB 1 RMB 2 RMB 50× cheaper
GPT-5.5 (International) ~210 RMB output / million tokens ~1/34 of DeepSeek V4-Pro output

This creates a fundamental economic truth: for applications with high cache hit rates (chatbots with system prompts, RAG systems, repetitive Agent tasks), the effective token cost is approaching zero. A system with 90% cache hits paying 0.025 RMB for those tokens and 3 RMB for the remaining 10% pays an effective blended rate of roughly 0.32 RMB per million input tokens — cheaper than the price of the electricity consumed to generate them.

Negotiation implication: When evaluating a Chinese AI vendor’s quote, the cache-hit rate is more important than the headline price. A vendor quoting 6 RMB/output but delivering 85% cache hits may be cheaper than a competitor quoting 2 RMB/output with 20% cache hits. Always request the vendor’s actual cache hit statistics for your workload class.

4. Time-of-Use Pricing: Tokens Like Electricity

In a landmark move, DeepSeek announced in 2026 that it would introduce time-of-use (TOU) pricing — the first major AI vendor globally to apply electricity-market logic to token pricing. Under the proposed scheme:

Time Period (Beijing Time) V4-Pro Input (Cache Hit) V4-Pro Input (Cache Miss) V4-Pro Output
Off-Peak (all other hours) 0.025 RMB 3 RMB 6 RMB
Peak: 9:00–12:00 & 14:00–18:00 0.05 RMB (2×) 6 RMB (2×) 12 RMB (2×)

This innovation reflects a deeper reality: token production is a spatio-temporal scarce resource. During peak hours, GPU clusters in western China run at maximum utilization; during off-peak hours, vast amounts of compute sit idle. By doubling prices during peak windows, DeepSeek converts idle capacity into incremental revenue while incentivizing customers to shift flexible workloads to off-peak hours — exactly as electricity markets have done for a century.

For foreign buyers, TOU pricing is both opportunity and risk. If your China-based AI workloads can tolerate scheduling flexibility (batch processing, overnight Agent runs, asynchronous inference), you can achieve effective token costs 50% below peak rates. If your workloads are real-time and user-facing, you must budget for peak pricing — and verify that your vendor’s quoted rate specifies which tier applies.

5. The International Price Gap & Export Arbitrage

The most strategically significant pricing fact for foreign firms is the persistent 10× to 34× gap between Chinese and international model output prices. The top five Chinese models’ output pricing stands at 1/10 to 1/34 of global giants like GPT-5.5.

1/10 – 1/34
Chinese output pricing vs. GPT-5.5
3–5×
Arbitrage when selling tokens to overseas customers
~1/60
DeepSeek V4-Pro output vs. Fable 5
~1/30
DeepSeek V4-Pro output vs. Claude Opus 5

This gap is not closing — it is widening. While Chinese vendors like Zhipu attempt to raise prices (and face ranking collapses), the leading players (DeepSeek, Doubao) are aggressively cutting prices to capture global developer mindshare. The economic logic is inexorable:

  • Cost advantage: Cheaper green power in western China, domestic chip economics (Huawei Ascend at 41% domestic share), and inference engine optimization combine to push Chinese token costs structurally below global equivalents.
  • Strategic intent: Chinese vendors are explicitly using low pricing as a tool for overseas expansion. The 3–5× international price arbitrage means a Chinese AI vendor can profitably sell tokens to European or Southeast Asian customers at prices that undercut local providers by 70–90%.
  • Risk for foreign partners: A vendor whose entire business model depends on unsustainable domestic pricing may not survive the next 12–18 months as the market consolidates. Conversely, a vendor with healthy overseas revenue (3–5× margins) has structural resilience.
Critical question for due diligence: Does your Chinese AI partner have a credible overseas revenue stream? If their entire business is domestic volume at razor-thin margins, their long-term viability is questionable. An independent professional enterprise credit report can verify their actual revenue composition, customer geographic distribution, and financial sustainability.

6. Contract Negotiation in the Token Era

For overseas procurement teams, the token pricing landscape of 2026 demands a fundamentally different negotiation approach than even 12 months ago. Here is a practical framework:

6.1 What to Watch in Every Quote

  • Cache-hit rate commitments: Request the vendor’s cache-hit statistics for workloads similar to yours. A vendor refusing to disclose this metric is hiding their true effective pricing.
  • Tiered pricing schedules: Verify whether quoted rates cover input (cache hit vs. miss separately), output, and any premium tiers. A single “per million token” number is meaningless without this breakdown.
  • Peak/off-peak differentials: Confirm which hours are classified as peak in the vendor’s TOU schedule and whether your workloads can be scheduled to avoid them. Get off-peak pricing locked in contractually.
  • Price escalation clauses: Given that Morgan Stanley data shows 48–80% price increases in just 12 months, contracts without price protection expose you to unilateral hikes. Negotiate annual caps or formula-based adjustments.
  • Currency and payment terms: RMB pricing exposes foreign buyers to FX risk. Explore whether the vendor offers USD/EUR denominated contracts with fixed exchange rate windows.

6.2 Your Negotiation Leverage

The market remains firmly a buyer’s market for foreign firms bringing meaningful volume. Vendors are still fighting for call volume to achieve scale economies. This means:

  • Multi-year committed use discounts: Lock in 12–24 month pricing with volume commitments. Vendors will trade price certainty for demand certainty.
  • Custom SLAs: Demand uptime guarantees, latency targets, and priority queue access — these are negotiable when you commit to volume.
  • Multi-vendor diversification: Use the threat of switching (DeepSeek ↔ Alibaba ↔ Huawei Cloud) as ongoing leverage. Vendors know lock-in is fragile in this market.
  • Capacity reservations: For mission-critical workloads, reserve dedicated capacity at fixed rates — this protects you from both price hikes and capacity shortages during peak demand.

6.3 Red Flags in Token Pricing

⚠️ Red Flag 1: Flat “Per Million Token” Quotes

Vendor quotes a single blended rate without separating cache-hit/cache-miss input and output. This obscures the true effective pricing and prevents you from optimizing your workload.

⚠️ Red Flag 2: No Price Protection

Contract allows unilateral price increases with minimal notice. In a market where prices swung 80%+ in 12 months, unprotected contracts are financially dangerous.

⚠️ Red Flag 3: Unsustainable Low Pricing

Vendor quotes prices below their own cost structure (verifiable through official enterprise credit reports). This is a cash-burn strategy that signals imminent price hikes or business failure.

⚠️ Red Flag 4: No Peak/Off-Peak Clarity

Vendor refuses to specify TOU schedules or won’t guarantee off-peak access. As DeepSeek’s model proves, peak pricing is the future — ambiguity here means budget uncertainty.

The strategic bottom line: Token pricing in China has moved from a simple “lower is better” equation to a sophisticated multi-dimensional optimization problem. The vendors who win are those who master cache architecture, chip adaptation, and spatio-temporal pricing. The foreign buyers who win are those who negotiate with full understanding of these dynamics — and who verify their partners’ financial resilience in a market transforming this fast.

🎯 Your 3-Step Action Plan for Token Pricing Negotiation

01

Verify before you sign. Commission an independent professional enterprise credit report on any Chinese AI vendor you are considering. Confirm their financial sustainability, revenue composition, and actual pricing model — because in a market transforming 80%+ per year, trust must be built on evidence.

02

Negotiate the effective rate, not the headline. Demand cache-hit statistics for your workload class, lock in off-peak pricing, secure multi-year price caps, and get peak-hour definitions in writing. The blended effective rate — not the marketing number — is what you pay.

03

Diversify and protect. Never depend on a single vendor. Design your China AI architecture for multi-vendor interoperability, reserve dedicated capacity for mission-critical workloads, and build contractual exit clauses. The K-shaped market will claim casualties — make sure you are not one of them.

At ChinaBizInsight, we help overseas firms navigate exactly these complexities. Our professional enterprise credit reports give you verified intelligence on Chinese AI companies’ pricing models, financial health, and business sustainability — because in a market where prices swing 80%+ annually, evidence is your only protection.

📚 References & Data Sources

DeepSeek official announcements (May 2026) — V4-Pro API permanently reduced to 25% of original pricing: cache-hit input 0.025 RMB, cache-miss input 3 RMB, output 6 RMB per million tokens; TOU pricing preview with 2× peak multipliers during 9:00–12:00 and 14:00–18:00 Beijing Time [3,9,12,18](@ref).

Tencent News (April 13, 2026) and Sina Finance (May 7, 2026) — Zhipu GLM-5 three price increases cumulative ~83%; OpenRouter ranking collapse from #3 (mid-Feb) to #17 (mid-Apr); Zhipu CEO Peng Zhang’s statement on “regressing token prices to normal commercial value” [5,14](@ref).

IMA Knowledge / industry analysis (July 2026) — DeepSeek-V2 May 2024 pricing at 1 RMB/1M input, 2 RMB/1M output (~1% of GPT-4 Turbo); Alibaba Tongyi Qwen-Long 97% price cut to 0.0005 RMB/thousand tokens; Doubao at 0.0008 RMB/thousand tokens (“99.3% cheaper than industry”) [1](@ref).

Morgan Stanley China AI Pricing Tracker (Q2 2026, reported August 2026) — Average Chinese API input price 4.9 RMB/million tokens (+48% vs. Q1 2025), output 21.9 RMB/million tokens (+80% vs. Q1 2025) [13](@ref).

Alibaba Cloud Developer Community (August 2026) — DeepSeek V4-Pro formal release: output at 6 RMB/million tokens = 1/60 of Fable 5 (~360 RMB), 1/30 of Claude Opus 5 (~180 RMB) [15](@ref).

Zhongshang Industrial Research Institute, China Token Factory Development White Paper 2026 — Chinese top-5 model output pricing at 1/10 to 1/34 of global giants; 3–5× arbitrage when selling tokens to overseas customers; cache-hit pricing approaching “milli-yuan” level with marginal cost near zero.

Note: All statistics current as of August 2026. DeepSeek’s announced TOU pricing is a “planned” mechanism as of August 2026; V4-Pro current formal pricing remains at the permanently reduced 25% level. Forward pricing is subject to vendor announcements and market conditions.

Your strategic bridge to transparent business in China.

Native Expertise
Direct Access
Official Sources
VIEW SAMPLES CONSULT EXPERT

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top