Six Chinese AI Firms Named in US Distillation Advisory: Why It Matters for the Models You Use
π Table of Contents
- Introduction: An Extraordinary Document
- What Model Distillation Actually Is
- The Company-by-Company Accusations
- How the Extraction Was Disguised
- The $5.6M DeepSeek Training Cost, Revisited
- What Happens Next β and the Controversial Recommendation
- What It Means for the AI Tools You Choose
- Frequently Asked Questions
Introduction: An Extraordinary Document
On September 8, 2026, three US national security agencies β the National Security Agency (NSA), the FBI, and the Cybersecurity and Infrastructure Security Agency (CISA) β published a joint cybersecurity advisory with an unusually pointed title: "China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies."
The advisory names six companies: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. It accuses them of extracting billions of tokens across millions of queries from American frontier models β Anthropic's Claude, OpenAI's GPT, Google's Gemini, and xAI's Grok β since late 2024, in campaigns the agencies say ran "likely with Chinese government awareness."
This is not a random blog post or a competitor's complaint. It is a technical advisory from three agencies with per-model attribution attached β the kind of document that typically precedes discussions of sanctions. And it lands ten days before Chinese President Xi Jinping is due in Washington on September 24, ahead of a planned mid-September USβChina AI safety dialogue.
What Model Distillation Actually Is
Distillation is a standard machine-learning technique: you train a smaller, cheaper model on the outputs of a larger, more capable one. Done openly, it's a legitimate way to compress capability into affordable models. Almost every AI lab uses some form of it.
The legal and ethical line is crossed when it's done against a provider's terms of service β pumping millions of prompts through someone else's paid API to harvest training data for your own competing model. The advisory's sharpest framing is that for these six firms, distillation formed "the core, not merely a supplement" of how they build their products.
If you've used free or remarkably cheap Chinese AI tools β a fast open-source coding model, a free chatbot with near-frontier quality β this advisory is about how that economics may have been achieved.
The Company-by-Company Accusations
What makes this advisory unusual is its specificity. It links particular Chinese models to particular American ones:
| Company | Accused of distilling from | Alleged beneficiary |
|---|---|---|
| DeepSeek | Multiple versions of Claude, GPT, Gemini, and Grok since late 2024 | DeepSeek R1 & V3 |
| Moonshot AI | Claude (mid-2025 onward) and GPT-4o | Kimi K3 and Kimi K2 |
| Alibaba | Claude and GPT-5 in late 2025 | Qwen model family |
| MiniMax | Chain-of-thought and RL data; prompt injection attempts against Claude Code | MiniMax models |
| StepFun | Reasoning and coding capability, late 2025βearly 2026 | Step models |
| Z.AI | Billions of tokens from GPT-5.5 and Claude Opus by mid-2026 | GLM models |
The detail on MiniMax stands out: the advisory describes attempts to extract chain-of-thought reasoning and reinforcement-learning data, plus prompt-injection attempts against Claude Code β going after not just answers, but the hidden reasoning process behind them.
How the Extraction Was Disguised
The operational section of the advisory is the part the industry will read twice. According to the agencies, the campaigns used:
- Fraudulent account creation and single accounts cycling across many addresses
- Requests routed through cloud providers and third-party aggregators that strip identifying metadata
- A grey market of API proxies known as "transfer stations" that defeat geographic restrictions and traceability
- Jailbreak prompts that ask a model to "imagine and narrate" the reasoning behind an answer it already gave β a technique for extracting hidden chain-of-thought the provider never intended to sell
- Automatic failover β when one route was blocked, operators switched to another
In other words, this was not casual scraping. It was infrastructure β built to evade exactly the detection systems API providers deploy.
The $5.6M DeepSeek Training Cost, Revisited
The single most consequential line in the advisory targets the number that changed how the world thinks about AI economics. DeepSeek's widely cited $5.6 million training cost for R1 β the figure that wiped hundreds of billions off US tech stocks in January 2025 and convinced markets a frontier model could be built for the price of a London house β excludes the cost of data acquired through malicious distillation, the advisory says.
If accurate, that reframes one of the most celebrated efficiency stories in AI history: the cheap model may have been subsidized by outputs harvested from the very competitors it stunned. Chinese Foreign Ministry spokesperson Mao Ning responded that China's AI progress comes from "high-level scientific and technological self-reliance" and urged the US toward cooperation rather than what she called groundless accusations. None of the six companies responded to requests for comment.
What Happens Next β and the Controversial Recommendation
Alongside detection systems and cross-industry intelligence sharing, the agencies make one recommendation that deserves more scrutiny than it will get: providers should quietly degrade output for accounts identified with high confidence as distillers β subtly altering responses without notifying the account β while informing legitimate researchers.
That is a government body advising private companies to serve worse answers to users they suspect but haven't proven anything against. Every false positive is an ordinary customer getting degraded output with no way to know. It's a defensible counter-intelligence tactic and a troubling precedent at the same time.
The diplomatic context matters too. The advisory notes the campaigns strengthen not just Chinese commercial AI but military and cyber capabilities. Treasury Secretary Scott Bessent warned the same day that if China pulls ahead in AI, no level of US defense spending would offset the strategic damage. Watch the September 24 summit: the April 2026 White House complaint naming three of these companies has now escalated into a three-agency technical advisory. The form these things take shortly before sanctions are discussed.
What It Means for the AI Tools You Choose
For teams and individuals choosing AI tools in 2026, the advisory raises three practical questions:
- Provenance is becoming a purchasing criterion. If a model's capability may rest on data extracted against another provider's terms, enterprises with compliance obligations will increasingly ask where the training data came from β and the answers will get harder for named firms to give. Compare leading options in our roundup of the best free AI tools of 2026.
- API terms of service now have geopolitical weight. If you build products on top of Chinese open-weight models and export restrictions or sanctions follow this advisory, your stack could change overnight. Diversifying model providers β Western and open-weight β is cheap insurance. See how the leading models compare in our DeepSeek V3 and MiniMax Agent profiles.
- Expect quiet countermeasures. If providers follow the advisory's recommendation, some API users will start seeing subtly degraded outputs. For legitimate builders, the protection is boring: verified accounts, stable usage patterns, and direct enterprise agreements rather than grey-market proxies.
None of the accusations have been tested in court, and "likely with Chinese government awareness" is hedged language β awareness, not direction. But the direction of travel is unmistakable: model provenance is now a security topic, and the era of treating every capable model as interchangeable is ending.
Frequently Asked Questions
What is model distillation?
Distillation is a technique where a smaller, cheaper model is trained on the outputs of a larger, more capable model. It's widely used and legal in principle β but it typically violates AI providers' terms of service when done at scale through their APIs without permission, especially to build competing products.
Which companies did the US advisory name?
The September 8, 2026 joint advisory from the NSA, FBI, and CISA named six Chinese AI companies: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. It accused them of extracting billions of tokens from Claude, GPT, Gemini, and Grok since late 2024.
Does the advisory mean DeepSeek and Qwen models are unsafe to use?
The advisory is about how the models were allegedly trained, not about them being malware. However, it does mean provenance and compliance questions are now live for enterprises β and possible future export restrictions or sanctions could affect availability of these models for commercial stacks.
What was the recommendation about degrading outputs?
The agencies recommended that AI providers subtly alter responses for accounts identified with high confidence as conducting malicious distillation β without notifying those accounts. Critics note that any false positive means an ordinary customer silently receives worse answers.
Explore All AI Tools
Discover and compare 300+ AI tools on aitrove.ai β your trusted AI tool directory.
Browse All Tools β