“Limits of State-Led Growth” Chinese-Language Data Famine Emerges as Next Bottleneck for China’s AI After Semiconductors
Authored On
Modified
Publicly available Chinese-language data accounts for just 1.3%, leaving AI training inputs in short supply Closed platforms and short-form dominance impede the accumulation of long-form knowledge Growing reliance on translation and synthetic data drives performance degradation and surging costs

China’s state-driven artificial intelligence (AI) push, which has rapidly expanded the country’s hardware footprint, has collided with a “Chinese-language data cliff.” Despite having 1.125 billion internet users and some of the world’s largest telecommunications and computing infrastructure, Chinese accounts for just 1.3% of all content across global websites. The convergence of state-controlled media, closed mobile- and video-centric platforms, and a shrinking open web has prevented sufficient accumulation of the long-form and specialized texts required to train large language models (LLMs). China’s AI industry is filling the void with translated English-language documents and synthetic data, but semantic distortion, model collapse, and copyright costs are increasingly constraining the pace of development.
High-Quality Public Data Projected to Run Out Within Six Years
According to the Hong Kong-based South China Morning Post (SCMP) on Aug. 10, Chinese AI experts and researchers have recently warned that the shortage of high-quality training data could become a new bottleneck for the country’s technological ambitions, eclipsing even hardware supply constraints. Epoch AI, a global research institute, projects that the supply of high-quality, publicly available text produced by humanity will be completely exhausted within the next six years.
Andrej Karpathy, a founding member of OpenAI, previously warned that a “data wall” could emerge within a decade and bring the growth of AI models to a standstill. Major U.S. technology companies are already pursuing aggressive measures, investing tens of millions of dollars to digitize offline books and scan printed volumes. Anthropic’s “Project Panama,” which involved cutting the bindings from millions of physical books before scanning them, illustrates the intensity of the competition to secure training data.
China’s data wall is considered particularly acute because of the linguistic characteristics unique to Chinese. According to web analytics company W3Techs, English accounts for nearly half, or 48.8%, of global web text, while Chinese represents just 1.3%, far below Spanish at 6%, German at 5.9%, and Japanese at 5%. The disparity is also evident in the latest figures from Common Crawl, which collects webpages directly. Chinese was the primary language in 4.43% of the HTML documents collected last month, compared with 40.58% for English. The figures indicate a severe shortage in the absolute volume of data required to train high-quality, Chinese-native AI models.
Chinese Once Expected to Overtake English, Yet Holds Just 1.3% of AI Training Data
In the early 2000s, forecasts proliferated that the rapid growth of China’s internet population would bring an end to English-language dominance. In 2011, when the number of Chinese-speaking internet users reached 400 million, observers predicted that Chinese would overtake English within five years. Those projections equated the number of users with the volume of content produced in each language.
China’s internet population subsequently grew exponentially. Data released by the China Internet Network Information Center (CNNIC) in February showed that the country had 1.125 billion internet users at the end of last year, with an internet penetration rate of 80.1%. China had 4.838 million 5G base stations, 42 large-scale AI computing clusters, and 1,590 exaFLOPS (EFLOPS) of intelligent computing capacity. The number of generative AI users had also risen to 602 million.
Connectivity and computing facilities have expanded to a world-leading scale, yet the production and public accumulation of language-based cultural assets in long-form and specialized texts have failed to keep pace. Growth in user numbers has not translated into a corresponding accumulation of searchable documents, specialized knowledge, and content preserved over extended periods. The value of high-quality data suitable for LLM training varies according to logical completeness, source traceability, terminological precision, and documented editing and revision histories. Telecommunications networks and data centers can be expanded rapidly through fiscal investment and standardized construction processes, while linguistic assets emerge from documents produced and validated over long periods by publishing, academic, journalistic, and archival ecosystems. The process requires sustained accumulation of materials preserving discipline-specific conceptual frameworks, changes in language across different eras, conflicting interpretations, and a wide range of cases.
Table 1. Principal Constraints on High-Quality Training Data for Chinese-Language LLMs
| Category | Key Conditions | Impact on AI Training Data |
|---|---|---|
| Mobile and video concentration | Smartphones are used by 99.6% of internet users. Short-form video users total 1.074 billion, or 95.4% of all users, while social media usage stands at 98.9%. | Weakens the foundation for accumulating searchable and reusable long-form text |
| Low knowledge density | Short videos and comments are heavily concentrated in repetitive expressions, buzzwords, and promotional language. | Widens the disparity between total data volume and high-quality training tokens |
| High processing costs | Video and audio require transcription, contextual classification, speaker identification, copyright clearance, and expert review of technical notation. | Increases the cost and processing time required to refine training data |
| Closed-platform ecosystem | WeChat Official Account articles and Douyin content accumulate inside their respective platforms, with access restricted for external search engines and web crawlers. | Constrains the acquisition and large-scale collection of publicly available data |
| Media controls | Authorities exercise broad control over the editorial direction, distribution scope, and continued availability of media and platform content. | Reduces argumentative depth and viewpoint diversity by suppressing content on Tiananmen, policy criticism, and public debate |
| Shrinking open web | The number of Chinese websites fell from 5.3 million in 2017 to 3.9 million in 2023. | The disappearance of independent websites, blogs, and online forums severs historical context and time-series data |
| Tighter generative AI regulation | Obligations cover data legality, intellectual property rights, consent for personal information use, accuracy, objectivity, diversity, and compliance with “core socialist values.” | Narrows the usability of legal, political, and modern-history materials while sharply increasing reclassification, verification, and safety-review costs |
Short-Form Overconsumption Deepens the “Knowledge Gap”
China’s internet ecosystem, however, remains disproportionately concentrated in smartphones and video content. According to CNNIC’s 57th Statistical Report on China’s Internet Development, 99.6% of internet users accessed the internet via smartphones, while short-form video users totaled 1.074 billion, representing 95.4% of the country’s internet population. Social media usage also reached 98.9%. With vast amounts of user time devoted to short videos and real-time communication, the foundation for accumulating searchable and reusable long-form text has weakened.
Video and audio can also serve as AI training resources, but they require separate processing procedures encompassing transcription, contextual classification, speaker identification, copyright clearance, and expert review of technical notation. Short videos and comments also contain high proportions of repetitive expressions, buzzwords, and promotional language, making it difficult to achieve sufficient knowledge density. This explains why the enormous volume of data generated through consumer activity does not translate directly into a comparable supply of high-quality tokens for LLM training.
Locked Inside Platforms and Erased by Censorship
The state-controlled media environment, in which content is edited and distributed under government oversight, also restricts the supply of publicly available data. Much of China’s internet content is consumed inside large applications. Documents published through official accounts on WeChat, China’s largest mobile messaging platform, and videos and comments posted on Douyin, the country’s leading short-form video platform, remain within their respective ecosystems, with access restricted for external search engines and conventional web crawlers.
This state-controlled media environment constrains the depth and diversity of Chinese-language content. The authorities exercise broad control over the editorial direction, distribution scope, and continued availability of media and platform content, suppressing long-form investigative reporting, policy criticism, and records of public debate from the production stage. Meanwhile, the central role of short videos, real-time posts, and officially guided commentary in information consumption has impeded the development of long-form content that preserves complex causal relationships and competing perspectives. Chinese-language materials available for LLM training consequently face quantitative scarcity alongside qualitative constraints, including weaker argumentative density, viewpoint bias, and fragmented time-series continuity.
The contraction of the open web is also undermining long-term accumulation. CNNIC statistics show that the number of websites in China fell from 5.3 million in 2017 to 3.9 million in 2023. Although the number of internet users rose rapidly over the same period, independent websites, blogs, and online forums were absorbed into mobile platforms or shut down. Server maintenance costs, migration to mobile platforms, content deletion, and censorship all contributed to the trend. As long-form records accumulated in newspaper articles, blogs, online forums, and personal websites disappeared, training materials preserving historical context and social debate vanished with them.
Regulatory barriers are also high. The Cyberspace Administration of China (CAC) implemented generative AI regulations in 2023, significantly tightening governance of training data. The rules placed responsibility on companies for ensuring the legality of data sources, protecting intellectual property rights, obtaining consent for the processing of personal information, and managing accuracy, objectivity, and diversity. Generated content was also required to comply with “core socialist values.” Materials acquired by AI companies must therefore undergo reclassification and verification under copyright, privacy, and political-compliance standards. In highly sensitive fields such as law, politics, and modern history, the range of usable training literature is narrowing, while the costs of data refinement and safety reviews are rising sharply.
Translation and Synthetic Data Fill the Void as Chinese-Language Sources Dry Up
China’s AI industry has responded with translation and synthetic data. The approach involves translating English-language research papers, books, and specialized documents into Chinese and feeding textbook-style passages and conversational materials generated by existing models back into training. Alibaba’s Qwen3, for example, was pretrained on 36 trillion tokens spanning 119 languages and dialects. Alibaba researchers reportedly extracted text from vast numbers of PDF files and used mathematics- and coding-specialized models to generate synthetic data, expanding the pool of training materials. The strategy reflects the assessment that Chinese-language source material collected from the open web cannot meet the required scale of data.
Chinese-centric models also rely substantially on English-language data. CT-LLM, a Chinese-focused model unveiled in 2024 by an international research team from the Hong Kong University of Science and Technology, Peking University, Fudan University, the University of Waterloo, Kuaishou, and other institutions, was trained on a total of 1.2 trillion tokens: 800 billion Chinese-language tokens, 300 billion English-language tokens, and 100 billion code tokens. The researchers separately constructed a curated dataset of Chinese-language web materials, AI-generated textbook-style synthetic data, and conversational materials. Translating English-language sources into Chinese enables rapid absorption of knowledge in science, technology, and medicine. Research presented last month at the Annual Meeting of the Association for Computational Linguistics (ACL) also found that translating knowledge-dense English documents into low-resource languages for pretraining improved performance on complex knowledge tasks.
Efforts to Offset Data Scarcity Undermine Model Performance
A growing share of translated data, however, can weaken linguistic naturalness and semantic precision. Awkward phrasing and inaccurate terminology introduced through machine translation can become entrenched as formulaic model output after large-scale pretraining. The risk of distortion is especially high in fields such as law, history, and philosophy, where institutional and cultural contexts carry substantial weight. The same term can differ across linguistic communities in semantic scope and the strength of the judgment it conveys, raising the likelihood of conceptual distortion and misinterpretation of causal relationships. Comparison against source texts, terminological standardization, and review by specialists in each field become indispensable, inevitably extending data-refinement timelines and increasing costs.
Synthetic data carries more complex risks. Repeatedly feeding AI-generated text into the training of successive model generations can eliminate low-frequency expressions and exceptional cases while amplifying existing errors. A study published in the international scientific journal Nature in 2024 confirmed that recursively training generative models on their own outputs can trigger “model collapse,” in which models lose the distribution of real-world data. Majority viewpoints and conventional sentences are particularly prone to overproduction, while rare knowledge and non-mainstream expressions risk being displaced during training.
Channels for directly collecting foreign web content are also narrowing. ByteDance deployed its web crawler, Bytespider, to collect documents extensively from English-language websites, but Cloudflare data show that its share of requests among major crawlers plunged from 14.1% in July 2024 to 2.4% in July 2025. Content providers have increasingly blocked AI crawlers and placed training data behind paywalls, making free collection more difficult. The strategy of compensating for the shortage of original Chinese-language material through greater volumes of translated content now faces both a performance ceiling and mounting verification costs.
- Previous “Energy, Logistics and Agriculture on High Alert” Europe Reels From Heat Waves and Drought as Need for Climate Adaptation Investment Mounts
- Next Southeast Asia’s Renewable Energy Push to Cut Middle East Dependence Stalls as Grid Bottlenecks Derail Half of Projects, Putting Data Center Drive at Risk