Since the beginning of 2025, the list of those accusing DeepSeek of “distillation” has grown longer and longer: OpenAI, White House officials, and Anthropic have all entered the fray, each offering numbers more specific than the last. But when these allegations are lined up, a counterintuitive fact emerges: after a full year, no party has produced evidence that can be reviewed by an independent third party; there has not been a single lawsuit; even OpenAI, which first raised the alarm, has publicly said it does not plan to sue. The precision of the accusations, and the absence of evidence, define the real shape of this controversy.
First, it is necessary to clarify a concept that is often glossed over. Distillation is a model compression technique, proposed in 2015 by Geoffrey Hinton and others, in which the outputs of a large model, the teacher, are used to train a smaller model, the student. It is widely used in both academia and industry and is entirely lawful in itself. The controversy has never been about the technique of “distillation” as such, but about a specific allegation: whether DeepSeek, without permission and in violation of terms of service, used outputs from OpenAI or Anthropic models to train its own models. This is a factual question about the origin of the inputs, not a question about the legitimacy of the technology.
Now lay out the “accusation space.”
In January 2025, DeepSeek surged in popularity because of R1, after which OpenAI, via the Financial Times, said there were signs that DeepSeek had used its GPT models through distillation; Microsoft was reportedly said to have detected suspicious API data outflows in the autumn of 2024. White House official David Sacks went further on television, saying there was “substantial evidence” — but that evidence has never been made public. In early February 2025, OpenAI’s CEO told reporters in Tokyo that there were no current plans to sue DeepSeek. One year later, in February 2026, OpenAI submitted a memorandum to the U.S. House Select Committee on China, accusing DeepSeek of “free-riding” on capabilities developed by OpenAI and other U.S. frontier labs, and saying it had detected “new, obfuscated methods.”
That same month, Anthropic entered the dispute and provided the most specific numbers to date. In its official report dated February 23, 2026, Anthropic accused three Chinese labs — DeepSeek, Moonshot, and MiniMax — of extracting Claude’s capabilities on an “industrial scale” through roughly 24,000 fraudulent accounts and more than 16 million interactions. The itemized figures were: more than 150,000 for DeepSeek, more than 3.4 million for Moonshot, and more than 13 million for MiniMax. The report said DeepSeek was the most technically sophisticated, using a method that induced Claude to gradually write out the internal reasoning behind an already completed answer — in effect, generating chain-of-thought training data at scale. Four months later, Anthropic, in a letter to the U.S. Senate, separately accused Alibaba of carrying out what it called the “largest known distillation attack to date,” involving roughly 25,000 accounts and 28.8 million interactions.
These allegations are certainly specific. But they share one common feature: all are unilateral claims. The three accused Chinese labs did not respond to requests for comment; Alibaba denied the accusation. No party — whether Anthropic or OpenAI — has made public any forensic evidence that an independent third party could review, such as a complete chain linking account ownership, or causal proof connecting the extracted content to capability gains in the target model. More importantly, there is one key counterpoint: OpenAI, the first to raise the alarm, explicitly said it would not sue. A plaintiff holding “substantial evidence” would not usually choose to forgo litigation, the forum where evidence can be placed before a court and subjected to cross-examination.
Now look at what the accused side has produced in the “evidence space.” The answer is: almost only one document, and one with a very narrow scope.
In September 2025, the DeepSeek-R1 paper, with Liang Wenfeng as corresponding author, appeared on the cover of Nature. In its 83 pages of supplementary materials, DeepSeek directly addressed the origins of its training data:
The training data for DeepSeek-V3-Base came only from ordinary web pages and e-books and contained no synthetic data; the authors also acknowledged that some web pages were observed to contain large amounts of answers generated by OpenAI models, which could have allowed the base model to indirectly acquire knowledge from other strong models; but during the pretraining cooldown phase, they “did not intentionally include” synthetic data generated by OpenAI.
This denial must be read precisely: first, it covers only the pretraining phase of V3-Base, and makes no denial whatsoever about whether outputs from third-party models were used in post-training, reinforcement learning, or SFT stages; second, the qualifier “did not intentionally include” leaves room for OpenAI-generated content passively mixed in from the web. This is a carefully worded denial, not a categorical statement that “we never used it.”
Thus the true shape of the controversy becomes visible: the accusations are highly specific, but all unilateral, without third-party review and without litigation; the denial does exist, but its scope is extremely narrow and its wording leaves room. The blank space in between has remained unfilled for a year. What makes that blank space especially hard to falsify is that DeepSeek still has not disclosed its full training dataset — it has open-sourced model weights and a large number of tool libraries, but the most critical data origins remain a black box. This does not constitute evidence that “it distilled,” but it does mean the proposition that “it did not distill” cannot be externally examined.
Some third-party research has also attempted to quantify the issue. A paper by the Shenzhen Institute of Advanced Technology of the Chinese Academy of Sciences, Tsinghua University, and other institutions, which has been accepted by ACL 2025, proposed a framework for measuring the “degree of distillation” in models. Its conclusion was that most well-known open-source and closed-source models exhibit a relatively high degree of distillation, with the exceptions of Claude, Doubao, and Gemini; base models show a higher degree of distillation than aligned models. This kind of research provides directional circumstantial support, but what it measures is similarity in model output distributions. It cannot identify “who distilled whom,” much less serve as conclusive proof of attribution.
This also leads to the hard counterpoint that must be made: the accusers’ own input side is also a glass house. Anthropic has just reached a $1.5 billion settlement over the use of pirated books as training data, in Bartz v. Anthropic, with final approval granted in July 2026; The New York Times’ copyright lawsuit against OpenAI is still ongoing. When these labs fed copyrighted texts into their models, they did not obtain permission from the rights holders either. The strongest counterquestion from the DeepSeek camp lies precisely here: if “using others’ outputs to train a model” should incur liability, should the same standard not also apply to the accusers themselves? A line circulating online and attributed to the DeepSeek side — “if it really was stolen, where could it have been stolen from?” — has no named firsthand citation, and the closest account is itself contradictory. It can only be treated as a circulating remark, not as DeepSeek’s official response.
Taken together, the honest conclusion is this: as of now, “DeepSeek distilled American models” remains a suspicion with nontrivial credibility, not a publishable fact. There is directional circumstantial support, but no conclusive attribution evidence, and zero judicial determination. The more important question in this controversy may not be “does distillation count as theft,” but rather: why do all the parties holding the data prefer to repeat accusations in the court of public opinion, rather than place the evidence in any forum where it can be independently examined?
Evidence List
Unilateral accusations: claims by one party, without third-party adjudication
- Anthropic report (2026-02-23): roughly 24,000 accounts and more than 16 million interactions across three companies; more than 150,000 for DeepSeek, more than 3.4 million for Moonshot, and more than 13 million for MiniMax; the three companies did not respond.
- Anthropic letter to the Senate (2026-06-10): roughly 25,000 Alibaba accounts and 28.8 million interactions, described as the “largest to date”; Alibaba denied it. The Alibaba figures and the figures for the three companies are two separate sets and should not be conflated.
- OpenAI memorandum to the House of Representatives (2026-02-12, obtained by Bloomberg): DeepSeek “free-riding” and “new obfuscated methods.”
- OpenAI in 2025-01, via the FT: “signs” of distillation.
Established: primary materials
- Altman on 2025-02-03 in Tokyo: no plans to sue DeepSeek.
- Nature R1 paper (DOI 10.1038/s41586-025-09422-z, Liang Wenfeng as corresponding author), supplementary materials: V3-Base used only web pages and e-books, with no synthetic data; synthetic data from OpenAI was “not intentionally included.” The scope is limited to V3-Base pretraining only.
- Distillation quantification paper (arXiv:2501.12619, ACL 2025): most open-source and closed-source models have relatively high distillation scores, except Claude/Doubao/Gemini.
- Bartz v. Anthropic: $1.5 billion settlement, final approval on 2026-07-20; the accuser’s input side as a glass house.
- David Sacks’ “substantial evidence” statement on Fox News: the evidence has never been made public and must not be written as an established fact.
- “If it really was stolen, where could it have been stolen from?”: this can only be included in the steelman as a counterquestion relayed by TMTPost, with no named firsthand citation, and with appropriate qualification.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接