ChinaBizInsight

CHINA AI INFRASTRUCTURE SERIES · PART 7 (FINAL)

The Utilisation Crisis: Why 30% GPU Utilisation Is China’s AI Infrastructure’s Biggest Hidden Risk

A thousand-card intelligent computing center in a western Chinese city has a server onboarding rate below 50% — and among those servers that are plugged in, actual GPU utilisation is below 30%. Annual operating cost: over 30 million RMB. This is not an outlier. According to the Institute of Artificial Intelligence at Inspur, China’s average intelligent computing center utilisation rate is just 30%. For every 100 units of compute capacity built across China, only about 30 are actually producing tokens. In an industry where capital expenditure runs into billions, this is the elephant in the room — and the single most important due diligence variable for any foreign firm partnering with Chinese AI infrastructure.

1. The Utilisation Numbers: A Tale of Four Benchmarks

The “30% problem” is not a single statistic — it is a cascade of overlapping metrics that tell a consistent story of systemic underuse. Understanding the distinctions is critical for accurate due diligence:

<30%
Industry-average GPU utilisation (Tencent Cloud, MWC Shanghai 2026; industry surveys)
30%
Average intelligent computing center utilisation (Inspur AI Institute)
20–30%
Rack-level average utilisation across deployed clusters
<10%
Enterprise self-built cluster utilisation (some 36Kr-surveyed cases)

The picture sharpens when we look at specific cases. Science and Technology Daily reported on a western Chinese city’s intelligent computing center: a thousand-card-scale facility with server onboarding rate below 50% and actual GPU utilisation below 30%, despite annual operating costs exceeding 30 million RMB. The investigation found that across China, “for every 100 units of compute capacity built, only 30 are actually used” is the norm — not the exception.

The problem is starkest with domestic chips. In a revealing on-site interview, Xinhua reporters found that in the same data center, servers equipped with NVIDIA GPUs had a leasing rate above 90%, while servers with domestic GPUs — though priced much lower — had a leasing rate below 50%. The reason, according to the operator: “The efficiency gap is huge, especially in the ecosystem.”

Cluster Type / Scenario Onboarding Rate Actual GPU Utilisation Key Driver
NVIDIA GPU clusters (prime locations) >90% High (constraint: supply, not demand) Mature ecosystem, developer familiarity
Domestic GPU clusters (same DC, adjacent racks) <50% Low-to-moderate Ecosystem immaturity, longer adaptation cycles
Western small-city intelligent computing centers <50% <30% Blind construction, weak local demand
Eastern enterprise self-built clusters N/A (owner-occupied) <10% (some surveyed cases) Sporadic usage, no multi-tenant sharing
Optimised clusters with scheduling platforms 60%+ achievable (from 30% baseline) Intelligent scheduling, PD disaggregation, peak/off-peak
Important statistical nuance: The Ministry of Industry and Information Technology reported a “national computing facilities overall utilisation rate of 71.4%” in August 2026 — but this measures facility-level energy and space utilisation, not GPU computational efficiency. The GPU “effective FLOPS utilisation” that determines token economics is a completely different — and far lower — metric. When a Chinese vendor quotes you a “utilisation rate,” always ask which definition they are using.

2. The Three Structural Mismatches

The utilisation crisis is not accidental — it is the predictable result of three deep structural mismatches in how China built its AI compute capacity. Engineer and scholar Zheng Weimin, academician of the Chinese Academy of Engineering, put it bluntly at WAIC 2026: “China’s AI industry does not have a global computing capacity shortage problem. The core pain point is concentrated in structural supply-demand mismatch of computing power.”

🎯

Mismatch 1: Training Supply vs. Inference Demand

The 2023–2024 capacity build-out was optimised for training peaks: high-bandwidth, all-to-all communication patterns requiring tightly coupled clusters. But post-2025, real growth is in inference — bursty, fragmented, latency-sensitive, and poorly suited to massive training clusters. Training clusters running inference workloads are like container ships used for river delivery: technically capable, economically disastrous. As one industry analysis notes, “simply adding more GPUs does not directly enhance productivity — the key lies in how to make existing computing power more efficient.”

🔗

Mismatch 2: General Supply vs. Specialised Demand

Some domestic chips carry immature software ecosystems; model adaptation cycles stretch to months. The fatal chain-break: “chips without software stack, software without model.” A domestic GPU cluster may be physically installed and powered on, but until major model providers port their architectures to the domestic software stack, those GPUs simply sit idle. Hence the 50% leasing gap between NVIDIA and domestic GPU racks in the same data center. Until the software ecosystem matures, domestic silicon is cheap hardware waiting for a compiler that can unleash it.

🌐

Mismatch 3: Regional Supply vs. Latency Demand

Western China offers the lowest cost (see Topic 3), but inference demand is concentrated in the east where users and applications live. “East Data West Computing” works beautifully for training (latency-tolerant, bandwidth-hungry) but breaks for inference. As one Shanghai engineer observed: “If an autonomous vehicle has to send its data to the west and back to Shanghai, the car will have driven several kilometres already.” Millisecond-level response requirements mean inference compute must stay close to users — yet most capacity was built in the west. The result: western clusters underused for inference, eastern capacity overcrowded and expensive.

The industry is responding with a new paradigm: “training moves west, inference sinks east.” Massive training clusters (ten-thousand to hundred-thousand card scales) are consolidating west of the Hu Huanyong Line, while inference-specific clusters are proliferating in the Yangtze River Delta, Greater Bay Area, and Beijing-Tianjin-Hebei region. Shanghai’s Lingang computing valley has already executed the world’s first production-environment cross-provincial AI inference migration — relocating a task to Hubei Shiyan in 3 minutes and reducing standalone load by 75%. This architectural rebalancing is essential, but it takes years — and many vendors will not survive the transition.

Compounding factor — fragmentation: Even within individual clusters, the Agent-era workload profile (input-output ratios as extreme as 154:1) causes severe GPU misallocation. When a single GPU is asked to both “ingest massive historical context” and “generate output tokens,” it oscillates between “starved and saturated” — creating computational bubbles and memory fragmentation that destroy effective utilisation. Software-layer innovations like “pipeline disaggregation” (specialising GPUs for either context ingestion or token generation) can dramatically improve cluster FLOPS efficiency — but require sophisticated runtime systems that most operators lack.

3. The Financial Consequences

For an industry that has invested hundreds of billions of RMB in compute infrastructure, the utilisation crisis is not an operational nuisance — it is an existential financial threat. The mathematics are unforgiving:

  • The fixed-cost trap: A ten-thousand-card cluster represents billions in CAPEX and hundreds of millions in annual fixed OPEX (power, cooling, real estate, depreciation). At 30% utilisation, 70% of that fixed cost is effectively spent producing nothing.
  • The 10-point rule: Every 10 percentage points of utilisation loss raises per-token cost proportionally. A cluster designed for 70% utilisation but running at 30% sees its per-token cost balloon by more than 130% versus plan — destroying any hope of profitability at prevailing token prices (see Topic 5).
  • The “busy but losing” paradox: Many 2024 projects remain unoperational — “half-finished” centres with powered shells but no tenants. Operational projects often rent raw compute at rock-bottom prices just to show activity metrics, creating a vicious cycle: low utilisation → low prices → negative margins → inability to invest in software/ecosystem → even lower utilisation.
  • The depreciation cliff: GPU hardware depreciates on a 3–5 year curve. A cluster that never reaches viable utilisation before its hardware generation becomes obsolete (accelerated by Huawei Ascend and Cambricon’s rapid iteration) faces total asset write-offs. China’s “15th Five-Year Plan” period (2026–2030) will see this cliff claim many victims.
The counter-example that proves the point: China Telecom’s Hangzhou “Xirang” platform, running a domestic TPU 1,024-card cluster, raised resource utilisation from 30% to over 60% through intelligent scheduling. The result: without adding a single piece of hardware, deliverable tokens nearly doubled and per-token comprehensive cost dropped significantly. This demonstrates that the utilisation gap is not a hardware problem — it is a software, operations, and business-model problem. The 30%→60% lever is worth billions.
Utilisation Scenario Effective Per-Token Cost Index Business Viability Typical Cluster Profile
70% (design target) 100 (baseline) Profitable at current prices NVIDIA clusters, prime locations, anchor tenants
50% (moderate underuse) ~140 Marginal — depends on pricing power Some domestic GPU clusters with software maturity
30% (industry average) ~233 Loss-making at prevailing token prices Western intelligent computing centers, mixed hardware
<10% (severe idle) >700 Existential — asset value impairment Enterprise self-built, fragmented demand

4. Due Diligence Questions That Matter

For overseas M&A teams, investment committees, and law firm due diligence practice groups, the utilisation crisis transforms token pricing from a commercial question into a survival question. Here is the specialised due diligence framework:

⚠️ Red Flag 1: “We don’t disclose utilisation”

Any Chinese AI infrastructure company that cannot or will not provide actual GPU utilisation statistics (by cluster, by month, by customer segment) is hiding a problem. Utilisation is the single most diagnostic metric in this industry. An independent professional enterprise credit report can often surface power consumption, equipment depreciation, and revenue-per-card ratios that indirectly reveal true utilisation.

⚠️ Red Flag 2: Capacity built ahead of anchored demand

If a vendor’s expansion story is “we built 10,000 cards and are now looking for tenants,” their utilisation will be disastrous. The correct order — as we explore in Section 5 — is “lock in token buyers first, then scale.” Vendors who built like real estate developers (supply-push) are the ones sitting at 20–30% utilisation with crushing depreciation.

⚠️ Red Flag 3: Domestic chip over-concentration

Clusters dominated by domestic GPUs (without proven model adaptation) face the 50% leasing gap. Verify which chips power which workloads. A vendor claiming “90%+ utilisation” on a purely domestic-chip cluster is either lying or counting facility-level metrics, not GPU FLOPS efficiency. The official enterprise credit report documents equipment procurement, supplier relationships, and depreciation schedules that reveal hardware composition.

⚠️ Red Flag 4: No scheduling platform, no peak/off-peak strategy

Vendors without intelligent scheduling software, peak/off-peak pricing (see Topic 5), or cross-region inference migration capability are structurally incapable of improving utilisation. They are trapped at 30%. Ask specifically: what is your PD-disaggregation strategy? What is your cross-region inference scheduling capability? What is your overseas token export plan (see Topic 6)? Silence on these questions is disqualifying.

Your due diligence checklist should include:

  • Demand the utilisation breakdown: By cluster, by month, by customer type (training vs. inference), by chip type. Aggregate “average utilisation” numbers are meaningless — you need the distribution.
  • Verify the anchor tenants: Who are the top 5 customers by token consumption? Are they locked in with multi-year contracts? A cluster with 3–5 committed anchor tenants at 70%+ utilisation is worth 10× a cluster with thousands of sporadic users at 25%.
  • Audit the depreciation schedule: How fast are GPUs being written down? A vendor depreciating over 5 years while the hardware becomes obsolete in 3 is inflating paper profits. Request the actual depreciation policy.
  • Assess the software stack: What scheduling platform? What model adaptations have been completed? What is the domestic-chip model porting pipeline? This determines whether utilisation can improve.
  • Map the geographic strategy: Training-west/inference-east rebalancing requires deliberate architectural planning. Is your vendor executing this transition, or are they stuck with west-only capacity while inference demand explodes in the east?
  • Review the financials holistically: An standard business credit report reveals revenue composition, power costs, depreciation expenses, and customer concentration — all of which indirectly quantify the utilisation reality that vendors may not disclose directly.

5. The “Build-to-Order” Solution

The root cause of the utilisation crisis is architectural: China built Token factories like real estate — construct first, find tenants later. This supply-push model worked in the early gold rush of 2023–2024 when any GPU was precious. It fails catastrophically in the mature market of 2026, where tokens are abundant and only effective tokens (low-cost, low-latency, high-quality) command value.

The industry is rapidly converging on the correct order:

The new sequence:
Lock in token buyers first — anchor tenants with committed volume contracts
Architect for their specific workload — training vs. inference mix, latency requirements, chip preferences
Scale capacity to matched demand — build incrementally against committed consumption
Layer on intelligent scheduling — peak/off-peak, caching, cross-region inference migration
Only then expand to token export (see Topic 6) with proven domestic utilisation as the foundation

“Production-based-on-demand” (按需生产) is rapidly becoming an unofficial threshold for new project approvals in the “15th Five-Year Plan” period. Provincial governments that once offered blanket subsidies for any compute construction are now conditioning support on demonstrated demand. Shanghai’s model — where “training moves west, inference sinks east” — is being replicated nationally: west-coast green-energy mega-clusters for committed training workloads, eastern edge inference clusters for milliseconds-latency applications.

The winners in this new paradigm share three characteristics:

  • Demand-pull, not supply-push: Capacity expands only against signed token purchase agreements. The vendor’s growth is gated by customer commitment, not by hardware availability.
  • Software-defined efficiency: Intelligent scheduling platforms that raise utilisation from 30% to 60%+ without adding hardware. This is where the real margin is created.
  • Diversified monetisation: Combining domestic token sales, overseas token export (Topic 6), and peak/off-peak arbitrage (Topic 5) to maximise revenue per installed card.
The investment implication: A Chinese AI infrastructure company at 60%+ utilisation with anchored demand and scheduling software is a fundamentally different asset than one at 30% utilisation with speculative capacity. The valuation gap is not 2× — it is 10×. Yet both may present themselves with impressive “total FLOPS” marketing numbers. This is precisely why token-level due diligence — not hardware-level due diligence — separates successful investments from catastrophic ones.
The meta-lesson across all 7 topics: China’s AI infrastructure story is not a simple “more compute = more value” narrative. It is a story of effective compute — tokens produced at low cost, low latency, and high reliability. From the western green-energy clusters (Topic 3) to domestic chip substitution (Topic 4), from the price war (Topic 5) to global token export (Topic 6), the unifying thread is utilisation efficiency. The vendors who master it will dominate. The vendors who ignore it will join the “half-finished centres” — powered shells with no future.

🎯 Your 3-Step Action Plan for Infrastructure Due Diligence

01

Quantify utilisation at the GPU level. Do not accept facility-level or “average” utilisation claims. Demand cluster-by-cluster, month-by-month, chip-by-chip breakdowns. Commission an independent professional enterprise credit report to verify power consumption, depreciation, and revenue-per-card ratios that indirectly reveal true GPU FLOPS efficiency.

02

Verify demand-pull architecture. Assess whether the vendor follows build-to-order discipline: anchored tenants, committed volume contracts, incremental capacity scaling. Avoid supply-push vendors who built speculative capacity ahead of demand — they are the ones drowning at 20–30% utilisation with accelerating depreciation.

03

Evaluate the efficiency roadmap. Does the vendor have intelligent scheduling software? A training-west/inference-east rebalancing strategy? Peak/off-peak pricing? Overseas token export capability? These are the levers that convert 30% utilisation into 60%+ — and they determine whether your investment compounds or craters. An official enterprise credit report documents the financial realities behind the vendor’s efficiency claims.

Across this 7-part series, we have explored China’s AI infrastructure from every angle — from the Token Factory model to the western green-energy clusters, from the chip substitution battle to the global token export frontier. Through it all, one truth has remained constant: in the token economy, effective utilisation is everything. At ChinaBizInsight, we help overseas firms separate the 60%-utilisation winners from the 30%-utilisation survivors. Our professional enterprise credit reports give you the verified intelligence you need to make infrastructure investment decisions with confidence — because in a market where 70% of compute sits idle, evidence is the only antidote to expensive mistakes.

📚 References & Data Sources

Science and Technology Daily (July 17, 2025) — “Utilisation Only 30%: How to Activate ‘Sleeping’ Computing Power”: Western city thousand-card intelligent computing center with server onboarding rate <50% and actual GPU utilisation <30%; annual operating cost >30 million RMB; Inspur AI Institute estimates China’s average intelligent computing center utilisation at 30%; as of November 2024, nearly 150 intelligent computing centers in operation nationwide, nearly 400 under construction or planning.

Xinhua Net (March 18, 2026) — On-site reporting: NVIDIA GPU servers in same data center leased at 90%+ rate, domestic GPU servers at <50% leasing rate; “efficiency gap is huge, especially in the ecosystem”; Feiteng Information VP Guo Yufeng confirms “many intelligent computing centers have computing power utilisation below 30%, large amounts of computing resources idle long-term”; industry structural problems of “valuing computing power,轻视 applications; valuing construction,轻视 effectiveness.”

China Energy Net (August 14, 2026) — “Training Moves West, Inference Sinks East”: Wang Chaoyang (Tsinghua) confirms ten-thousand to hundred-thousand card clusters have basically migrated west of the Hu Huanyong Line; eastern regions will see massive inference cluster proliferation; UCloud founder Ji Xinhua notes even 70–80% utilisation at edge still leaves fragmented idle compute.

People’s Daily / Xinhua (August 17–18, 2026) — MIIT reports national computing facilities overall utilisation rate of 71.4%; intelligent computing capacity reached 2,185 EFLOPS by end of June 2026 (177% YoY). Note: This 71.4% is facility-level utilisation, distinct from GPU effective FLOPS utilisation of 30%.

China Science and Technology Network (July 17, 2025) — Eastern region thousand-card intelligent computing center with onboarding rate <50% and actual utilisation <40%; enterprise self-built clusters average utilisation <30%, some only 10%+; actual cost expenditure 10× higher than planned 100%-utilisation assumption.

Economic Information Daily / China Economic News (June 26, 2026) — Tencent Cloud insights: “Continuously improving resource utilisation” is the core solution; Tencent 2025 capital expenditure 79.2 billion RMB, 2026 Q1 capex 31.94 billion RMB (+16% YoY), constrained by GPU supply; “we want to buy cards, but for a long time we couldn’t get them.”

Central Broadcasting System (July 23, 2026) — WAIC 2026: China Telecom Hangzhou “Xirang” platform raised cluster resource utilisation from 30% to 60%+ for a 1,024-chip domestic TPU cluster; utilisation doubling means “without hardware investment, deliverable tokens nearly double, single-token comprehensive cost drops significantly.”

Pengpai News (March 7, 2026) — “15th Five-Year Plan” computing network strategy: “training west migration, inference east sink” decoupling model; Shanghai Lingang computing valley executed world-first production-environment cross-provincial AI inference migration (3-minute relocation to Hubei Shiyan, standalone load -75%); China Telecom “Xirang” platform offers hourly billing, single domestic AI card rental cost down 60%.

Interface News (August 19, 2026) — Agent workloads show 154:1 input-output ratio causing GPU misallocation; software/runtime layer “pipeline disaggregation” (PD disaggregation) assigns context-ingestion GPUs vs. token-generation GPUs to eliminate memory fragmentation and computational bubbles; cross-region, cross-time-slot scheduling evolution from “where there are cards, send there” to “what task, what time, should go where.”

People’s Daily (Shanghai) (August 13, 2026) — Domestic computing clusters moving toward “usable + profitable” via co-creation ecosystems.

Note: All statistics current as of August 2026. GPU effective FLOPS utilisation (30% industry average) and facility-level utilisation (71.4% per MIIT) are distinct metrics and should not be conflated in due diligence assessments. The “build-to-order” model is an emerging industry norm for the “15th Five-Year Plan” period and is not yet uniformly regulated.

Your strategic bridge to transparent business in China.

Native Expertise
Direct Access
Official Sources
VIEW SAMPLES CONSULT EXPERT

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top