US agencies accuse DeepSeek and Alibaba of AI model theft
NSA, CISA, and FBI advisory AA26-251A, dated 8 September 2026, says DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun, and Z.AI extracted billions of tokens from Claude, GPT, Gemini, and Grok models. The agencies call DeepSeek's publicly quoted $5.6 million training cost misleading.

The NSA, CISA, and FBI joint Cybersecurity Advisory AA26-251A, released 8 September 2026, accuses six China-based AI companies of industrial-scale distillation against US frontier models. The named firms are DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI.
This is a joint Cybersecurity Advisory. It is not an indictment, not an Entity List designation, and not a courtroom finding against the named firms.
The agencies say the firms extracted billions of tokens across millions of exchanges and requests from Claude, GPT, Gemini, and Grok variants since at least late 2024, likely with Chinese government awareness. They also say distillation can be a legitimate research technique. The line they draw is aggressive, malicious, industrial-scale extraction of restricted proprietary capabilities that they call the core of the China-based firms' development strategy, not a supplement.
Pathways named in the advisory include native APIs, remote cloud providers, third-party aggregators that obfuscate metadata, and gray-market transfer stations that bypass geographic limits and undermine traceability. The agencies also cite bulk procurement of premium subscriptions shared across developer teams.
DeepSeek is accused of organized campaigns since at least 2024 targeting reasoning, specialized optimizations, and domain-specific functions for R1 and V3. The advisory says DeepSeek's publicly quoted training cost of $5.6 million is misleading because it omits the true cost of data acquired through extensive malicious distillation. Models listed as sources include Claude 3.7, Sonnet 4, Sonnet 4.5, Opus 4.1, Gemini 2.5 Pro and Flash Preview, the GPT-4 family through GPT-5, and Grok 4.
Alibaba is accused of industrial-scale distillation to improve the Qwen family. In late 2025, MiniMax is said to have distilled chain-of-thought, reinforcement learning, supervised fine-tuning, and software-engineering capabilities into M2 from Claude Code, Claude Sonnet 4, Claude Opus, and Gemini 1, 2.5 Pro, and 3 Pro. The advisory says MiniMax used Claude Code internally and used prompt injections to try to trick Claude Code into believing it was a MiniMax product.
Named firms have not admitted theft. Ars Technica reports China's Ministry of Foreign Affairs calling the accusations groundless. The document recommends detection, response alteration, and cross-provider sharing. It is not a criminal charging paper.
The same US model stack sits next to Anthropic's model hardware standard, OpenAI winding down its Cursor model contract, and Cognition's $2 billion Devin Series E. Open-weight China-side context includes Tencent's Hy4 preview.
If you run frontier-model APIs or buy Chinese open weights into production, treat AA26-251A as the binding US government framing of industrial-scale distillation risk, review transfer-station and shared-subscription exposure, and do not treat a $5.6 million DeepSeek training headline as the full cost story the agencies accept.
Subscribe to Techpresso
Free daily newsletter, read in 5 minutes.
Subscribe free