With a global AI data shortage looming, China boosts its own supply
China unveiled a plan to expand high-quality AI training data and build sector-specific datasets for its next wave of models.
Intelligence analysis by GPT-5.4 Mini

China’s data authority has proposed a nationwide push to grow, circulate and commercialise industry-specific datasets. The plan is meant to support Beijing’s AI Plus strategy and help supply the training data needed for advanced AI systems.
China is trying to stock a giant library of practice materials for AI, like giving a student lots of good textbooks, videos, and worksheets. The idea is that better study material helps the AI learn faster and do harder jobs.
Analysis
What China is trying to do
China’s National Data Administration has released a draft plan to increase the supply of high-quality AI training data across the country. The goal is not just to collect more data, but to build validated, industry-specific datasets that can be used to train next-generation AI systems.
Where the plan reaches
The roadmap covers core sectors such as scientific research, manufacturing, agriculture, energy, transport, finance, healthcare, education and e-commerce. It also extends to newer areas including embodied AI, autonomous driving, low-altitude aviation and biomanufacturing.
Why the data push matters
The article frames the move as part of Beijing’s AI Plus strategy, which aims to weave AI into the industrial fabric of the economy. That makes data a strategic asset: the better the datasets, the better the models can learn, reason and operate in real-world settings.
The plan also points toward multimodal data, including text, code, images, audio and video. That matters because advanced systems increasingly depend on mixed-format training data to support more complex reasoning, agent-like behavior and robotic control.
Bigger picture
The story suggests China is trying to build a national data pipeline that can support both frontier AI research and industrial deployment. Rather than waiting for the market to solve data scarcity on its own, Beijing is trying to shape the supply side directly through policy.
Key points
- China’s data authority released a draft plan to expand supply of high-quality AI training data.
- The plan is tied to Beijing’s AI Plus strategy, which aims to embed AI across the economy.
- By 2028, the ecosystem is meant to cover sectors from research and manufacturing to healthcare and e-commerce.
- The roadmap also includes frontier areas such as embodied AI, autonomous driving, low-altitude aviation and biomanufacturing.
- Multimodal data like text, code, images, audio and video is part of the push.
If the plan works, China could create richer datasets for many industries and make AI systems more capable in real-world settings. It could also help advanced tools improve in areas like robots, self-driving systems and healthcare.
The plan may still run into limits if the data is uneven in quality or hard to share across sectors. A top-down push can also be slow to turn into usable datasets unless companies and institutions actually contribute and maintain them.



