U.S. Agencies: China Firms Extracted Billions of AI Tokens
NSA, CISA and FBI say China-based AI firms pulled billions of tokens from U.S. frontier models since late 2024 to train rival systems.
The National Security Agency, Cybersecurity and Infrastructure Security Agency and FBI reported that China-based AI companies extracted billions of tokens from U.S. frontier models beginning in late 2024 and used those outputs to train competing systems.
The joint report names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI and says the firms queried variants of Anthropic’s Claude, OpenAI’s GPT series, Google’s Gemini and Meta’s Grok across millions of exchanges. Between late 2024 and mid-2025, the document details DeepSeek drawing capabilities from Claude, Gemini, GPT-4 and GPT-5 and Grok 4 to train its R1 and V3 models. It describes Moonshot using Claude Fable 5 outputs to improve its Kimi-K3 model and GPT-4o outputs to enhance Kimi-K2.
Analysts mapped the tactics, techniques and procedures used in the campaigns to the MITRE ATLAS framework. The mapping covers stages from resource development and access through collection and exfiltration. The report identifies the kinds of knowledge distilled into rival models, including API rule-driven task behavior, agentic functions, question-and-answer optimization, supervised fine-tuning techniques and improvements in creative and occupational writing.
The agencies also identified methods not cataloged in MITRE ATLAS. These include regional restriction evasion, subscription exploitation, centralized request routing infrastructure, automated request metadata sanitization and systematic quota and cost optimization. The report characterizes these practices as planned and well-resourced national-level activity and states the campaigns were “likely with Chinese government awareness.”
On impact, the document records financial harm to U.S. model developers and calls the activity a “strategic economic threat to fair technological competition and U.S. technological leadership.” The agencies noted that knowledge about model defenses obtained through extraction could enable more sophisticated attacks, while stressing the primary harm described is to industry competitiveness and leadership.
To limit harm, the report recommends coordinated mitigations across cloud providers, API aggregators and infrastructure firms. Suggested defensive steps include behavioral detection and monitoring to flag high-volume or patterned requests that indicate automated distillation. The agencies advise applying differential privacy by adding calibrated noise to model outputs to reduce the risk that outputs reveal training membership information or signals that could reconstruct private data.
The document also proposes more assertive responses where confidence is high, such as targeted changes to impose costs on malicious distillation campaigns, and sharing multi-source correlated intelligence to improve attribution and justify response actions with lower risk to legitimate users.
Background material in the report explains knowledge distillation as repeatedly querying a target model and using its outputs to train another system. The agencies provide examples of request patterns that can signal malicious distillation and encourage vendors to adopt detection tools and privacy-preserving techniques to reduce large-scale extraction risk.








