Anthropic Reports Record Distillation Campaigns by Chinese AI Labs
Anthropic releases a report Thursday alleging persistent and escalating distillation attacks by China-based AI companies aimed at harvesting the capabilities of its Claude models. The company attributes the activity to five separate campaigns totaling nearly 200 million exchanges, making the efforts larger and more aggressive than earlier incidents it describes in February. The attackers target some of Claude's most valuable capabilities, including agentic tool use, coding, data analysis, and logical reasoning.
Distillation attacks work by extracting a model's chain of thought from its responses, which can then train smaller models through supervised fine-tuning. Anthropic normally hides its models' internal reasoning behind summarized thinking blocks, but the campaigns discover techniques that trick Claude into revealing its thinking traces directly. In one example, an attacker frames a request as a translation task, asking the model to translate its working memory into katakana-only Japanese.
The largest campaign, attributed to Alibaba, accounts for 151 million exchanges between May and July 2026, peaking at nearly three million exchanges per day. Anthropic links the activity to a single effort because the 3,500 accounts involved share one fixed prompt for extracting chain of thought, and it appears designed to produce training material for Alibaba's Qwen models. A separate campaign from Moonshot AI, maker of Kimi, seems to route requests directly from the Chinese military, according to the report.