China's race to build next-generation AI models is entering a new phase, and the threat is no longer just chips. Chinese AI experts increasingly warn that running out of high-quality training data could become the next major bottleneck for the nation's tech ambitions — one that hardware workarounds cannot easily solve.

The concern is global. According to Epoch AI, a US-based research institute, the global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years. OpenAI co-founder Andrej Karpathy has warned of a looming 'data wall' by the end of this decade, beyond which model capabilities could plateau unless fed fresh, reliable information.

The problem is sharper for China. Chinese makes up only about 1.3% of web content, compared with roughly 49% for English, yet Chinese-language data accounts for more than 60% of the training data used by most Chinese AI models. As models grow, demand for quality Chinese text is outpacing supply.

Beijing is already responding: authorities plan national datasets by 2028, according to reports, and Chinese labs are turning to synthetic data and multilingual sources. Meanwhile, top American labs are spending lavishly to mine offline human knowledge — a scramble that has ignited a fierce ethical debate about how far companies should go to feed their models.

The data crunch follows a year in which Chinese open-weight models such as Moonshot's Kimi K3, Alibaba's Qwen3.8-Max and DeepSeek's V4-Flash have closed much of the capability gap with US systems — and it threatens to slow that momentum at exactly the moment the country is pushing to lead the field.