On the afternoon of September 15, at the AI Security Governance Sub-forum of the National Cybersecurity Promotion Week 2026 held in Jinan, the Chinese Internet Basic Corpus 4.0 was officially released to the public and was fully launched on the "Chinese Internet Corpus Resource Platform." The total data volume of this release reached 120 GB, aiming to provide a solid and reliable data support for the high-quality development of large model training and the artificial intelligence industry in China.
Under the guidance of the Cyberspace Administration of China, the China Cybersecurity Association, in collaboration with the National Internet Emergency Response Center, and together with multiple industry units such as Baidu, Zhongke Wengge, Kepu Cloud, Torex, iFLYTEK, and Zhihu, jointly advanced this project. Building upon the previous successful releases of Chinese Internet Basic Corpus 1.0, 2.0, and 3.0, all parties have leveraged the corpus co-construction and sharing mechanism established by the Cyberspace Administration's AI Security Governance Committee to gather a batch of new high-quality trusted data. After a series of rigorous and detailed data processing measures, the 4.0 version was finally formed for public access.
A relevant official from the Cybersecurity Association pointed out that the official launch of the Chinese Internet Basic Corpus 4.0 is an important阶段性 achievement of society's collaborative efforts to build high-quality Chinese corpora, further enriching the domestic supply system of high-quality Chinese corpora. In the next step, the association will continue to work with the National Internet Emergency Response Center and other units, and coordinate with various industries to further deepen the construction of Chinese Internet Basic Corpus, laying a solid data foundation for the high-quality development of the entire artificial intelligence industry.
Users who need to obtain data can access the "Chinese Internet Corpus Resource Platform" through the official website of the China Cybersecurity Association, follow the registration and certification procedures, and then download or contact to obtain specific corpus data.


