Those Who Leave the Door Open for You to Distill, and Those Who Weld It Shut: The Overlooked Asymmetry in the DeepSeek Controversy

This article examines the overlooked asymmetry in the DeepSeek controversy: DeepSeek open-sources its models under the MIT license while its accusers keep theirs closed, raising the question of whether output-side openness can offset alleged input-side borrowing—and arguing that generosity and innocence are two separate ledgers.
Those Who Leave the Door Open for You to Distill, and Those Who Weld It Shut: The Overlooked Asymmetry in the DeepSeek Controversy

There is an asymmetry in this controversy that almost no one has mentioned. The companies accusing DeepSeek have closed-source models—you cannot even measure who they might have distilled from. DeepSeek, meanwhile, has open-sourced its model weights, its tool libraries, and now its entire agent framework under the most permissive MIT license—it openly invites the whole world to distill from it. So the question is no longer a simple "did they steal or not," but a far harder one: when a company lays all of its capabilities open, how do you weigh that alleged borrowing on the input side?

This asymmetry is concrete and verifiable, not a matter of sentiment. DeepSeek Harness was open-sourced under the MIT license in August 2026—one of the most permissive open-source licenses. Anyone can freely fork it, redistribute it, use it commercially, or even distill its behavior into their own systems. This is not an isolated case but a consistent posture from DeepSeek: it open-sources model weights, open-sources a large number of tool libraries, and is one of the most open frontier laboratories today. By contrast, OpenAI and Anthropic, the ones making the accusations, keep their core model weights locked tight. To use a blunt but accurate comparison: DeepSeek leaves the door open for you to distill; they weld the door shut to keep you out.

This raises a genuine question worth laying out seriously: on the moral ledger, can extreme openness on the output side offset the alleged borrowing on the input side?

Those who support DeepSeek would argue as follows. First, the induction and distillation of knowledge can hardly be called anyone's private property. When these American laboratories fed articles from The New York Times and Forbes into their models back then, they did not obtain consent from the rights holders; now that their own models' outputs are being learned by others, they appeal to "theft"—that is a double standard. Second, DeepSeek's technical narrative is internally consistent: its R1 paper attributes the source of its capabilities to architectural innovations such as GRPO reinforcement learning and MoE sparse experts (published in Nature and peer-reviewed), and it explicitly states that its post-training phase distills from its own DeepSeek-R1—that is "self-distillation," not distillation of a third-party model. Third, on commercial logic, according to observations by some analytical institutions, DeepSeek's official channels barely monetize inference services, which weakens the "stealing for profit" framework of accusation: a laboratory that does not intend to make money from model services and open-sources all of its results does not have a complete motivational chain for "stealing to sell."

If this article stopped here, it would be a soft piece defending DeepSeek. But for a series to hold up, it is precisely in the most sympathetic installment that the sharpest rebuttal must be saved for last.

Openness on the output side cannot automatically clear the provenance of the input side. There is a plain ethical intuition here: if the raw materials of a work were obtained without permission, then whether it is subsequently open-sourced or freely shared does not change the nature of the raw materials themselves. If a person makes bread from stolen flour and gives it away to neighbors for free, the bread is a good deed, but the flour is still stolen—the two cannot offset each other. Treating "I open-sourced everything" as moral immunity for an input-side problem is logically untenable.

Going further, "I open-sourced everything" may itself be a carefully cultivated narrative. The openness is real, and it is also strategically valuable: it wins developers, wins reputation, and wins the moral position of "victim-benefactor." When a company faces questions on the input side that it can hardly prove or disprove, redirecting public attention to its irreproachable generosity on the output side is an effective rhetorical diversion. Pointing this out is not to deny the genuine contribution of DeepSeek's open-sourcing, but to remind: generosity and innocence are two independent ledgers, and a surplus on one ledger cannot be applied to offset the doubts on the other.

There is one more detail that must not be blurred into ambiguity. The "self-distillation" written in DeepSeek's paper (distilling from its own R1) and the alleged "distilling third-party models" are two entirely different things. The former is a fully legitimate internal technical method; the latter is where the controversy lies. Any argument that substitutes "post-training used distillation" for "therefore it distilled OpenAI" is dishonest—this line must be held when speaking for DeepSeek, and equally held when criticizing it.

So this asymmetry ultimately leaves us not with an answer, but a more sober question. DeepSeek's openness is real and valuable—it makes the word "steal" far more awkward to use than it is against a closed-source company. But openness is not an indulgence for the input side. This series is willing to spend an entire installment loading one side of the scale for DeepSeek, precisely to prove: we are not witch-hunting; we are weighing. And honesty in weighing requires us to place weights on both ends—even if one of those ends is our own previous criticism.