On September 8, 2026, the US National Security Agency (NSA), the Cybersecurity and Infrastructure Security Agency (CISA), and the Federal Bureau of Investigation (FBI) jointly released cybersecurity advisory AA26-251A, calling out six Chinese AI companies—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—for allegedly using billions of tokens and millions of queries since late 2024 to systematically extract capabilities from US frontier AI models. The targets included Anthropic's Claude (multiple versions), OpenAI's ChatGPT (five versions), Google Gemini (two versions), and xAI's Grok 4.
The advisory uses terms such as "systematic," "industrial-grade," and "malicious" to describe the conduct, and explicitly asserts: "The knowledge distillation campaign constitutes the core of these companies' AI development strategy, not merely a supplement."
How a Legitimate Technique Became "Theft"
Understanding this controversy requires first clarifying what "knowledge distillation" itself is. The technique was introduced by Hinton's team in 2015, with the core logic of letting smaller models learn the output distribution of larger models, achieving capabilities close to those of large models at a far lower computational cost. Every major AI company globally, including OpenAI, Google, and Anthropic themselves, widely uses distillation internally to produce lightweight deployment versions. Distillation has never been considered illegal—until it is used to extract outputs from others' commercial APIs to train competing models.
This is precisely where the legal gray area of the accusation lies. The terms of service of OpenAI, Anthropic, and Google all explicitly prohibit "using API outputs to train models that compete with us," but these are contractual terms, not legal prohibitions. Violating terms of service means accounts can be banned, but it does not rise to criminal charges. US intelligence agencies choosing to characterize this as a "cybersecurity incident" and jointly issuing an advisory is a proactive move to elevate a civil contract dispute into a national security framework, rather than a forced response to a technological breakthrough.
The characterization itself is the core of the controversy.
Scale Is the Decisive Factor
The US agencies' argument is not that distillation itself is illegal, but rather emphasizes the anomaly of scale. According to the CISA advisory, the volume of queries used by Chinese companies constitutes "industrial-grade" operations: bypassing geographic restrictions and terms of service through fraudulent accounts, bulk premium subscriptions, and proxy routing services; dispersing requests across different accounts, platforms, and third-party aggregators to obfuscate metadata; and evading origin detection through relay station services.
The DeepSeek case detailed in the advisory is particularly specific: the company used four versions of Claude, two versions of Gemini, five versions of ChatGPT, and Grok 4 to generate synthetic training data for its R1 and V3 models, covering multiple capability dimensions including agent functions, question-answering optimization, and creative writing. The Moonshot AI case goes further—according to the advisory, training for Kimi K2 and K3 used outputs from a total of 18 US models, including Anthropic's most advanced commercial model, Fable 5.
This number is the key to understanding the severity of the accusation: 18 models means this is not an incidental API call but a systematic capability extraction project. Moonshot AI declined to comment on the matter.
The Very Existence of Kimi K3 Is Counter-Evidence
US agencies accuse Moonshot AI of distilling Fable 5 to train Kimi K3, yet after its release, Kimi K3 surpassed Fable 5 on the Frontend Code Arena benchmark. According to UK tech media outlet Tom's Hardware, Kimi K3 is an open-weight model with 2.8 trillion parameters, making it one of the largest open-weight AI models ever released.
This creates a paradox: if distillation were really just copying homework, why would the distilled model outperform the original on certain benchmarks? This anomaly has led multiple AI researchers worldwide to question some of the US agencies' technical judgments. According to the South China Morning Post, many AI experts believe that fully attributing Kimi K3's strong performance to distillation may underestimate Moonshot AI's own engineering capabilities.
Logs of large-scale API calls are auditable, and the US agencies' accusations are clearly based on observable behavior patterns. The key question—to what extent distillation is the true source of Kimi K3's capabilities—still lacks any public independent verification.
The Timing Makes This Advisory More Complicated
The US agencies chose to release this advisory on September 8, a highly sensitive moment. According to Reuters, US President Trump planned to invite Chinese President Xi Jinping to visit the White House in late September, with AI expected to be among the core topics of the talks. As recently as April, on the eve of Trump's visit to Beijing, the US government made similar accusations. Both accusations emerged on the eve of diplomatic summits, making it difficult to completely rule out the political logic of negotiation leverage.
Chinese Foreign Ministry spokesperson Mao Ning responded by stating that "China's AI development is an achievement of high-level scientific and technological self-reliance and self-improvement," and called on "the US side to earnestly implement the important consensus reached by the two heads of state and stop unfounded accusations and smears." This is a standard diplomatic denial—it neither directly addressed the specific accusations nor provided any rebuttal evidence.
In July, China made a counter-accusation of its own, claiming that US AI companies had used Chinese data samples to train their models. This strategy did not receive an equivalent volume of international media coverage, but it logically forms a mirror image: if a one-way data flow constitutes theft, then with a two-way data flow, who acted first?
The Countermeasures Recommended by US Agencies Are Unexpected
The part of this advisory most worth reading closely is not the accusation itself, but the response recommendations offered by the US agencies. The advisory recommends that US AI companies quietly degrade the response quality of accounts identified with high-confidence malicious distillation behavior, rather than banning them outright.
This is a highly pragmatic but controversial strategy. Direct bans would expose detection logic and force the other side to adjust its evasion techniques; degrading responses, by contrast, can continuously undermine the quality of their training data without triggering confrontation. But it also means major US AI platforms would provide different quality of service to specific accounts—creating tension with the open and neutral principles these companies profess.
More importantly, this recommendation presupposes a premise: that US AI companies have the technical capability to identify malicious distillation accounts. If the detection logic is insufficiently precise, large numbers of legitimate academic or commercial users would suffer silent degradation. This is not a hypothetical risk but a real predicament with abundant precedents in the content moderation field.
The Military Capability Link Is the True Core Objective
The key argument through which the US agencies place this accusation within the national security framework appears only later in the advisory: distillation "not only reduces the R&D costs of Chinese AI companies, but also enhances the military and cyberattack capabilities that China can use against the US and its allies." According to Reuters, Chinese military researchers have previously used outputs from US frontier AI models to train domestic military AI systems.
This is the real strategic destination of the advisory. Directly linking commercial API usage to military capability enhancement provides legal grounds for tougher regulatory measures—whether cutting off Chinese companies' access to US APIs or pushing allies to establish similar joint blocking mechanisms. The advisory explicitly calls for "a coordinated response across the AI ecosystem, including effective information sharing among the US government, private enterprises, and allied nations."
This is a narrative leap from bilateral commercial dispute to multilateral technology alliance mobilization.
Independent Assessment
This accusation can be evaluated separately on several levels.
At the factual level: large-scale calls to competitor APIs to extract training data most likely did occur. The scale of proxy routing and account dispersion exceeds any reasonable explanation for normal commercial use and corroborates existing independent reports. Anthropic confirmed similar distillation attack observations to CNBC as early as February 2026.
At the characterization level: framing commercial conduct that violates terms of service as malicious cyber activity and incorporating it into a national security advisory is a deliberate escalation. This escalation serves specific policy objectives and does not mean the technical facts themselves require a framework of this intensity to be described.
At the competition logic level: the fundamental reason distillation is used on such a large scale by Chinese companies is the continued US restrictions on GPU exports. Under conditions of constrained computing power, using knowledge distillation to compress computational costs is an entirely rational engineering alternative. US intelligence agencies calling this behavior the core of China's AI industrial policy—that judgment may be accurate, but it describes an adaptive response to restrictions, not an independent choice rooted in malice.
When the speed at which AI capabilities spread exceeds any single country's ability to control them, where will replacing standardization with confrontation lead the entire industry? At present, no party has offered a credible answer to this question.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接