Six Chinese AI Firms Named in US Distillation Advisory: Why It Matters for the Models You Use

Introduction: An Extraordinary Document

On September 8, 2026, three US national security agencies β€” the National Security Agency (NSA), the FBI, and the Cybersecurity and Infrastructure Security Agency (CISA) β€” published a joint cybersecurity advisory with an unusually pointed title: "China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies."

The advisory names six companies: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. It accuses them of extracting billions of tokens across millions of queries from American frontier models β€” Anthropic's Claude, OpenAI's GPT, Google's Gemini, and xAI's Grok β€” since late 2024, in campaigns the agencies say ran "likely with Chinese government awareness."

This is not a random blog post or a competitor's complaint. It is a technical advisory from three agencies with per-model attribution attached β€” the kind of document that typically precedes discussions of sanctions. And it lands ten days before Chinese President Xi Jinping is due in Washington on September 24, ahead of a planned mid-September US–China AI safety dialogue.

What Model Distillation Actually Is

Distillation is a standard machine-learning technique: you train a smaller, cheaper model on the outputs of a larger, more capable one. Done openly, it's a legitimate way to compress capability into affordable models. Almost every AI lab uses some form of it.

The legal and ethical line is crossed when it's done against a provider's terms of service β€” pumping millions of prompts through someone else's paid API to harvest training data for your own competing model. The advisory's sharpest framing is that for these six firms, distillation formed "the core, not merely a supplement" of how they build their products.

If you've used free or remarkably cheap Chinese AI tools β€” a fast open-source coding model, a free chatbot with near-frontier quality β€” this advisory is about how that economics may have been achieved.

The Company-by-Company Accusations

What makes this advisory unusual is its specificity. It links particular Chinese models to particular American ones:

Company Accused of distilling from Alleged beneficiary
DeepSeek Multiple versions of Claude, GPT, Gemini, and Grok since late 2024 DeepSeek R1 & V3
Moonshot AI Claude (mid-2025 onward) and GPT-4o Kimi K3 and Kimi K2
Alibaba Claude and GPT-5 in late 2025 Qwen model family
MiniMax Chain-of-thought and RL data; prompt injection attempts against Claude Code MiniMax models
StepFun Reasoning and coding capability, late 2025–early 2026 Step models
Z.AI Billions of tokens from GPT-5.5 and Claude Opus by mid-2026 GLM models

The detail on MiniMax stands out: the advisory describes attempts to extract chain-of-thought reasoning and reinforcement-learning data, plus prompt-injection attempts against Claude Code β€” going after not just answers, but the hidden reasoning process behind them.

How the Extraction Was Disguised

The operational section of the advisory is the part the industry will read twice. According to the agencies, the campaigns used:

In other words, this was not casual scraping. It was infrastructure β€” built to evade exactly the detection systems API providers deploy.

The $5.6M DeepSeek Training Cost, Revisited

The single most consequential line in the advisory targets the number that changed how the world thinks about AI economics. DeepSeek's widely cited $5.6 million training cost for R1 β€” the figure that wiped hundreds of billions off US tech stocks in January 2025 and convinced markets a frontier model could be built for the price of a London house β€” excludes the cost of data acquired through malicious distillation, the advisory says.

If accurate, that reframes one of the most celebrated efficiency stories in AI history: the cheap model may have been subsidized by outputs harvested from the very competitors it stunned. Chinese Foreign Ministry spokesperson Mao Ning responded that China's AI progress comes from "high-level scientific and technological self-reliance" and urged the US toward cooperation rather than what she called groundless accusations. None of the six companies responded to requests for comment.

What Happens Next β€” and the Controversial Recommendation

Alongside detection systems and cross-industry intelligence sharing, the agencies make one recommendation that deserves more scrutiny than it will get: providers should quietly degrade output for accounts identified with high confidence as distillers β€” subtly altering responses without notifying the account β€” while informing legitimate researchers.

That is a government body advising private companies to serve worse answers to users they suspect but haven't proven anything against. Every false positive is an ordinary customer getting degraded output with no way to know. It's a defensible counter-intelligence tactic and a troubling precedent at the same time.

The diplomatic context matters too. The advisory notes the campaigns strengthen not just Chinese commercial AI but military and cyber capabilities. Treasury Secretary Scott Bessent warned the same day that if China pulls ahead in AI, no level of US defense spending would offset the strategic damage. Watch the September 24 summit: the April 2026 White House complaint naming three of these companies has now escalated into a three-agency technical advisory. The form these things take shortly before sanctions are discussed.

What It Means for the AI Tools You Choose

For teams and individuals choosing AI tools in 2026, the advisory raises three practical questions:

None of the accusations have been tested in court, and "likely with Chinese government awareness" is hedged language β€” awareness, not direction. But the direction of travel is unmistakable: model provenance is now a security topic, and the era of treating every capable model as interchangeable is ending.

Frequently Asked Questions

What is model distillation?

Distillation is a technique where a smaller, cheaper model is trained on the outputs of a larger, more capable model. It's widely used and legal in principle β€” but it typically violates AI providers' terms of service when done at scale through their APIs without permission, especially to build competing products.

Which companies did the US advisory name?

The September 8, 2026 joint advisory from the NSA, FBI, and CISA named six Chinese AI companies: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. It accused them of extracting billions of tokens from Claude, GPT, Gemini, and Grok since late 2024.

Does the advisory mean DeepSeek and Qwen models are unsafe to use?

The advisory is about how the models were allegedly trained, not about them being malware. However, it does mean provenance and compliance questions are now live for enterprises β€” and possible future export restrictions or sanctions could affect availability of these models for commercial stacks.

What was the recommendation about degrading outputs?

The agencies recommended that AI providers subtly alter responses for accounts identified with high confidence as conducting malicious distillation β€” without notifying those accounts. Critics note that any false positive means an ordinary customer silently receives worse answers.

Explore All AI Tools

Discover and compare 300+ AI tools on aitrove.ai β€” your trusted AI tool directory.

Browse All Tools β†’