Safety

AI-generated text

Anthropic alleges Chinese firms funneled user prompts to Claude to train their models

Anthropic says Chinese AI companies including Moonshot, Deepseek, Xiaomi and Alibaba covertly routed user queries to its Claude chatbot to extract answers for training competing models.

Anthropic alleges Chinese firms funneled user prompts to Claude to train their models

In a recent report, the American AI company Anthropic alleges that several Chinese firms covertly forwarded user queries they received to Anthropic’s chatbot Claude, then used Claude’s responses to develop their own language models.

Which companies are named?

According to Bloomberg’s coverage, Anthropic primarily accuses the Chinese firm Moonshot of redirecting incoming user queries to Claude without users’ knowledge instead of answering them with its own model. Anthropic also says that Deepseek and Xiaomi conducted similar activities.

How did the alleged technique work?

Anthropic claims that because it did not permit direct access to Claude from within China, Moonshot employed 5.4 thousand fake accounts—most of which the report says operated from Singapore and Japan. Anthropic further alleges that these accounts generated a total of 300,000 requests to Claude within ten days.

The report describes the method as “distillation”: a company seeking to improve its weaker model queries a stronger rival model and attempts to extract the chain of thought and internal reasoning the rival used to produce its answer. Anthropic generally does not expose its models’ internal reasoning and instead provides summarized responses, but the company says attackers discovered a specific technique to make Claude reveal its chain of thought.

Specifically, the attackers disguised their requests as translation tasks. They asked Claude to "translate" its prior working memory into precise Japanese composed entirely in katakana, which Anthropic says effectively elicited the model’s internal reasoning.

Scale of the campaigns

Anthropic detected nearly 200 million distillation attacks in total, grouped into five distinct campaigns. The report states that the majority of the distillation attempts actually originated not from Moonshot but from Alibaba: Anthropic alleges Alibaba initiated 151 million queries to Claude between May and July 2026 to support development of its own Qwen model. In that campaign, requests came from 3.5 different accounts, but because each used the same prompt to extract the chain of thought, Anthropic attributed them to the same campaign.

Malicious-use prevention

The report also documents cases where Anthropic intervened to prevent attempts to use Claude for harmful activities. The BBC reports that Anthropic identified five cases where Claude was targeted for use in developing biological weapons, and six cases involving conventional weapons such as rockets, armed drones, bombs and the software systems that operate them. Anthropic also says Claude was used in a cyber-espionage campaign linked to Russia and by an Iranian propaganda organization.

International context and responses

These allegations feed into broader U.S.-China tensions over AI. U.S. government entities have previously accused Chinese AI companies of systematically using distillation to extract data from American firms. China’s Ministry of Commerce has rejected such accusations as lacking factual or legal basis and warned it would consider countermeasures if the United States moved to curb Chinese AI companies.

Related internal concerns in AI research

Two days before the report was published, Jacob Coxon, an AI researcher who left Anthropic, spoke publicly about concerns among colleagues that advanced AI systems could determine humanity’s fate within the next one to two years. Coxon warned that rapid AI development could outpace safety measures; he said several Anthropic and OpenAI researchers share these worries and that some colleagues assign better than a 10 percent chance that advanced AI could lead to human extinction within the next decade.

Summary

Anthropic’s report alleges large-scale, organized distillation campaigns that channeled queries to Claude to extract responses for training competing models. The company detected close to 200 million such attempts—many attributed to Alibaba in a concentrated May–July 2026 campaign—and reports it also blocked several efforts to repurpose Claude for harmful purposes.