China faces new AI bottleneck as it runs out of Chinese-language training data

China is facing a significant bottleneck in AI development due to a shortage of high-quality Chinese-language training data. Experts warn that this data scarcity could hinder the nation's tech ambitions, mirroring similar challenges faced by US companies.
Why it matters
The race for AI dominance is shifting from hardware to data acquisition, which has profound implications for global technological competitiveness and ethical standards in data mining.
China’s high-stakes race to build next-generation artificial intelligence models is entering a critical new phase, where a less visible yet far more existential threat is coming into view: a severe shortage of high-quality training data.
While the US chokehold on advanced computing chips has dominated headlines, Chinese AI experts increasingly warn that running out of quality data could prove to be the next major bottleneck to the nation’s tech ambitions – and one that hardware workarounds cannot easily solve.
It is a challenge confronting AI giants on both sides of the Pacific – and some US companies are already resorting to aggressive measures to stay ahead.
The global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years, according to US-based research institute Epoch AI.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in