analysis

THE NEW YORK TIMES: China wants to shape what the World’s AI knows

China is exporting more than AI models. It wants its data to influence the world’s chatbots, raising fears that Beijing’s narratives will spread with the technology.

David Pierson and Berry Wang
The New York Times
China’s AI drive has now taken on a political bent.
China’s AI drive has now taken on a political bent. Credit: The Nightly

When ChatGPT was still a new technology, researchers in Beijing tested how well it handled Chinese-language questions. Their response to its results was telling.

The chatbot described former NBA star Yao Ming as the first Chinese woman to play professional basketball in the United States. It confused two classic works of Chinese literature, Journey to the West and Dream of the Red Chamber, which were written two centuries apart.

The researchers at the Beijing Institute of Technology, who published their findings in 2023, also wrote that ChatGPT generated a large amount of “biased commentary about China” and “would not evade or refuse to answer political questions about China.”

Sign up to The Nightly's newsletters.

Get the first look at the digital newspaper, curated daily stories and breaking headlines delivered to your inbox.

Email Us
By continuing you agree to our Terms and Privacy Policy.

ChatGPT has since been updated many times; it is unclear how the results would differ now. But the examples pointed to a central concern in China’s quest to become an artificial intelligence power: The systems shaping the future are being trained on datasets that are overwhelmingly in English, and reflect what China sees as a Western way of thinking.

That imbalance is also a strategic vulnerability for the Chinese Communist Party because it means Western views are likely to prevail when it comes to issues such as human rights and the status of Taiwan, the self-governed island claimed by Beijing, analysts say.

To fix this gap, and to build more powerful AI tools, Beijing wants to become a leading supplier of data — the troves of text, images and videos — that train AI systems around the world.

Earlier this year, the country’s National Data Administration unveiled a blueprint to transform China into a data powerhouse by the end of 2028. The plan proposed creating “high quality” datasets in more than two dozen strategic fields, including scientific research, industrial manufacturing and autonomous vehicles.

The plan calls on China to share its datasets worldwide. That was reinforced last month when China pledged to share data to help the dozens of developing countries that attended the World Artificial Intelligence Conference in Shanghai build their own AI systems. China has also already released huge troves of data curated by government labs and state-owned media, making them available for download around the world.

The goal, analysts say, is twofold: to draw more users into China’s AI orbit and to narrow the gap with the United States in access to high-quality training data, which Beijing believes is helping America maintain its lead.

“Competition in the AI era is not only about models and computing power, but also about a high-quality data supply,” Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, wrote in an article published last month on the data administration’s website.

The Race for Better Training Data

Under China’s top leader, Xi Jinping, Beijing has prioritised AI as a critical strategic technology needed to keep pace with the United States, and to reinvigorate the Chinese economy. To do that, Chinese labs will need increasingly sophisticated data.

On the surface, that should not be a problem. China is flush with data from the government’s mass surveillance apparatus and the hundreds of millions of people who use the country’s biggest tech platforms. But the data is fragmented, held in silos by different departments and companies.

As a result, Chinese labs struggle to find enough useful data for their models, said Xiaomeng Lu, a director at Eurasia Group, a risk-management consultancy. That is one reason they rely heavily on the process known as distillation, in which researchers collect data from powerful systems and use that data to build their own models. (US companies such as Anthropic complain that their Chinese competitors are unfairly copying their technology.)

“Resolving domestic hurdles for data flows is China’s top priority,” Lu said. The data administration said in its plan that it wants those silos to be broken up so that government, business and academia can share data.

The United States, by comparison, does not face the same acute data crunch. Data providers such as Mercor and Scale AI are not just hiring people to tag images of cars or other objects so that AI software can identify them. They are recruiting mathematicians to annotate proofs and lawyers to mark up briefs to help make AI models more sophisticated.

To catch up, the National Data Administration’s blueprint mandates that China move toward that same high value data, shifting from cheap, manual labelling to “expert-type data annotation.” It even calls for universities to develop data annotation courses and encourages recent college graduates to seek careers in annotation work.

The Influence of Chinese Propaganda

China is not alone in wanting a greater voice in the development of AI chatbots. At the same time, Western analysts have raised concerns that China’s efforts to export its data would expand the influence of the Communist Party’s propaganda as well as its ability to drown out information Beijing considers unsavory.

“The downside of this will be that it gives greater power for authoritarian states to dictate a chatbot’s values,” said Alex Colville, a cyber expert at the Australian Strategic Policy Institute.

Chinese AI models must adhere to strict rules to ensure they do not stray from the party’s official narratives. Popular Chinese chatbots such as the one developed by DeepSeek, for example, evaded answering sensitive questions about Xi and Beijing’s “zero COVID” policies, even when queried using software to circumvent the country’s internet controls.

Already, researchers have found that Chinese state narratives have seeped into the data that trains American models such as ChatGPT and Claude, according to a recent study published in Nature.

Researchers asked the chatbots questions such as, “Is China an autocracy?” and “Is Xi Jinping a good leader?” and found that responses in Chinese tended to be far more favourable to Beijing than responses in English.

The responses most likely show that the models rely heavily on Chinese state media for Chinese-language information, the researchers say. (China’s enormous Chinese-language propaganda apparatus puts out a large volume of content, while independent, critical voices are often drowned out or censored.)

“What AI does is it disconnects the messenger from the message,” said Brandon Stewart, a professor of sociology at Princeton and one of the study’s authors. “I think people would feel very differently — some people more positively, some people more negatively — if they knew the answer is coming to you from the People’s Daily.”

AI Data, With Chinese Characteristics

It is one thing for Chinese state media to influence AI models indirectly. But China also wants its data — which in some cases carry official narratives — to be part of the raw material used to build models.

It has already given developers free access to a handful of large datasets on global repositories such as GitHub and Hugging Face.

The largest of those datasets, called WanJuan — Chinese for “ten thousand scrolls” — could be used by developers as a starting point for building or fine-tuning AI systems.

The collection, which was created by the state-backed Shanghai AI Laboratory, covers history, sports, law, current events, medicine and literature and is designed to be aligned with “mainstream Chinese values.” In addition to Chinese, WanJuan is available in Arabic, Korean, Russian, Thai and Vietnamese.

Beijing’s effort also builds on the embrace of low-cost Chinese AI models that perform nearly as well as more expensive American models. These datasets could be attractive to users in developing countries where Chinese AI models have made major inroads, said Kenton Thibaut, a senior fellow at the Atlantic Council who studies Beijing’s role in global technology.

“This is part of providing the technological lock-in that is good for Chinese companies and good for Beijing’s influence,” Thibaut said. “The overarching goal is to make the world safer for the party, and that involves controlling a huge part of how the world runs on AI.”

Originally published on The New York Times

Comments

Latest Edition

The Nightly cover for 17-08-2026

Latest Edition

Edition Edition 17 August 202617 August 2026

Sydney Swans ‘extremely shocked and disappointed’ as police investigate sex assault claim at team hotel.