2025-10-14
China Has Built Over 35,000 High-Quality Datasets with a Total Capacity of More Than 400 PB
Source:People’s Posts and Telecommunications News (RENMIN YOUDIAN)

  Recently, at the “Thematic High-Quality Dataset Exchange Event” of the 2025 China International Big Data Industry Expo, the High-Quality Dataset Construction Guidelines (hereinafter referred to as the Construction Guidelines) were officially released. According to the data, China has already built more than 35,000 high-quality datasets, with a total capacity exceeding 400 PB.

  The Construction Guidelines indicate that with the rapid development of large-model technologies, the focus of artificial intelligence (AI) research is shifting from “optimizing model architecture” to “co-optimizing models and data”, highlighting the increasingly prominent role of high-quality data. As one of the three core elements of AI development, data has become a fundamental resource for training large AI models and determines their performance. Accelerating the construction of high-quality AI datasets and consolidating the data foundation for AI development is of significant importance for promoting the implementation of “AI+” scenarios.

  In December 2024, the National Development and Reform Commission, the National Data Administration, and other departments issued the Guiding Opinions on Promoting the High-quality Development of the Data Industry, which, for the first time, clearly defined the concept of “high-quality datasets” and positioned them as a core carrier for integrating AI with the real economy. Subsequently, a series of policies were introduced, including the Implementation Opinions on Promoting the High-quality Development of the Data Annotation Industry, the Opinions on Promoting the Development and Utilization of Enterprise Data Resources, and the Guidelines for the Development of National Data Infrastructure, all emphasizing the construction of industry-specific high-quality datasets.

  Under these policy directives, China has achieved significant results in high-quality dataset construction. Data published in the Development Guidelines show that by June 2025, more than 35,000 high-quality datasets had been built nationwide with a capacity totaling over 400 PB; 3,364 high-quality datasets have been listed by data trading institutions, serving as key commodities in circulation, with cumulative transaction volumes reaching nearly RMB 4 billion and a total scale of 246 PB; in domestic model training, Chinese-language data now accounts for 60%–80% of usage.

  The National Data Administration has coordinated the construction of data labeling bases, piloting initiatives in ecosystem development, capability enhancement, and scenario applications. By gathering leading enterprises, it has promoted regional AI industry ecosystems. To date, 524 industry-specific high-quality datasets have been built, totaling over 29 PB of data, supporting the development and application of 163 domestic AI large models and generating related output value of over 8.3 billion RMB in the data labeling industry. Meanwhile, central enterprises, large-model technology companies, standardization organizations, and academic research institutions are jointly building an industry ecosystem, forming a multi-faceted, coordinated development pattern.

  The Development Guidelines note that although China has unique advantages in national coordination, advancement models, and application scenarios, there remain gaps in data openness, standard systems, key technologies, and international influence. Challenges also persist in data supply, technical tools, standardization, security and compliance, and business models.

  The Development Guidelines emphasize the need to optimize the construction layout of high-quality datasets through systematic thinking, facilitate dataset circulation and utilization through infrastructure, ensure sustainable development in an ecological environment, and establish a high-quality dataset framework covering all processes and links. It advocates building an industry knowledge indexing framework to meet intelligent demands, mapping industry dataset resources for targeted smart scenarios, and constructing a comprehensive, industry-wide standard system for the construction and operation of high-quality datasets.

  At the same time, by establishing integrated “platform + dataset + model” service facilities, the aim is to lower the barriers for dataset application and promote market circulation and large-scale utilization. Through institutional innovation, industry collaboration, and talent cultivation, a multi-party win-win ecosystem can be built for tackling bottlenecks such as high construction costs, low willingness to share, and weak innovation momentum.