US DOJ Steps In to Back OpenAI: AI Training Constitutes Fair Use, but a Key Conflict of Interest Is Sidestepped

On September 1, 2026, the US Department of Justice filed a statement of interest in The New York Times' copyright lawsuit against OpenAI and Microsoft, arguing that training large language models on copyrighted works constitutes fair use—while remaining silent on reported negotiations for the federal government to acquire a stake in OpenAI.

On September 1, 2026, the US Department of Justice (DOJ) filed a "statement of interest" with the Manhattan federal court, formally intervening in The New York Times' copyright infringement lawsuit against OpenAI and Microsoft. The DOJ explicitly argued that using copyrighted written works to train large language models (LLMs) constitutes "fair use" under US copyright law and should not incur infringement liability. The filing was co-signed by Deputy Attorney General Stanley Woodward Jr. and others.

This marks the federal government's first formal position on the series of AI copyright lawsuits, covering The New York Times case as well as related actions brought by book authors and publishers being heard jointly in the same period. Presiding Judge Sidney Stein had previously required both sides to submit motions for summary judgment by September 4 at the latest.

What the Training Process Actually Does in Legal Terms

"Fair use" under copyright law is assessed by four factors, the most critical being "the purpose and character of the use" (factor one) and "the effect of the use upon the potential market for the original work" (factor four). The DOJ's argument centers on these: the training process does not present a New York Times article verbatim to readers, but rather converts all text into numerical representations, allowing the model to learn statistical relationships among vocabulary, syntax, and knowledge, thereby acquiring capabilities such as text prediction, editing, translation, and generating new content.

Training is an "exceedingly transformative" act because the model never uses the article for the author's original purpose of conveying information or providing entertainment to readers. On factor four, the DOJ argued that the training act itself does not substitute for the original articles or harm their market value—readers will not stop subscribing to The New York Times because ChatGPT exists. The DOJ specifically noted that what must be evaluated separately is the model's output stage: if a particular output reproduces protected content verbatim and distributes it to users, that is a separate legal question, but it cannot be used to retroactively render the training act itself unlawful.

The DOJ clearly delineated boundaries in the filing: the statement pertains only to the training stage, covering neither the methods of acquiring and storing training data nor model outputs that may reproduce protected content. This means that even if the court adopts the government's position, OpenAI's liability risk at the output level remains.

The National Security Card, and an Undisclosed Detail

The DOJ placed part of its argumentative weight in this filing on national security issues. The filing cited an executive order previously signed by Trump—"Removing Barriers to American Leadership in Artificial Intelligence"—and explicitly stated: "Legal rules that make it significantly more difficult to build a strong AI industry in the United States threaten national security and confer competitive advantages on foreign adversaries not subject to such constraints."

If every training run required obtaining licenses from copyright holders beforehand, compliance costs would rise sharply; the DOJ noted that only the largest technology companies could ultimately absorb these costs, which would instead intensify LLM market concentration and grant traditional media organizations with large historical archives an asymmetric advantage.

Just as the DOJ was filing this document in support of OpenAI, the Trump administration and OpenAI were reportedly in negotiations for the federal government to acquire approximately 5% equity in OpenAI—worth about $42.6 billion at OpenAI's then-valuation of roughly $852 billion. The entire statement of interest made no mention of this whatsoever. According to Above the Law, this potential financial relationship has prompted questions within the legal community about a conflict of interest.

DOJ intervention in litigation is not unusual, but it is typically grounded in public policy interests rather than commercial ones. Expressing a position to the court on a defendant's core legal claims while potentially holding equity in that same defendant is a highly sensitive backdrop in legal proceedings, and the absence of any disclosure in the filing makes the issue all the more conspicuous.

Winners and Losers Among Stakeholders

For AI developers, the symbolic value of the government's statement far exceeds its legal binding force. A "statement of interest" itself carries no adjudicative effect, but it sends a clear signal to the judge presiding over the case: the executive branch believes that imposing copyright liability on training practices would have consequences extending beyond a single case. In the US legal system, a government statement of interest typically exerts some influence on the court, particularly on matters involving national policy.

For content creators and media organizations, the situation is decidedly unfavorable. A New York Times spokesperson, Graham James, responded that the government "is siding with a handful of trillion-dollar AI companies at the expense of countless American creators." He added: "AI companies simply need to pay fairly for the content used in their products, as copyright law requires." Steven Lieberman, lead attorney for the New York Daily News, noted that the government's position "ignores the Copyright Clause of the US Constitution" and contradicts views previously expressed by the US Copyright Office.

For small and medium-sized AI startups, the government's argument cuts both ways: on the one hand, if training is deemed fair use, data compliance costs fall and market entry barriers lower; on the other hand, the DOJ itself acknowledged that if training required obtaining licenses case by case, the ultimate beneficiaries would be precisely the "largest technology companies"—meaning that a license-free regime does not necessarily alter market concentration, as leading companies' advantages in compute, data scale, and engineering capability would continue to dominate the industry landscape.

For downstream developers and enterprise users, the most immediate short-term impact is reduced legal uncertainty. Pending a definitive court ruling, the signal that "the government considers training to be fair use" at least partially alleviates compliance anxieties surrounding the use of large model APIs.

Historical Precedents and What Makes This Case Different

The DOJ cited two persuasive precedents in its filing. The first is Google v. Oracle, in which the Supreme Court ruled in 2021 that Google's copying of Java API code constituted fair use, with the core rationale being the transformative nature of the use; the second is the Second Circuit's ruling in Authors Guild v. Google, which held that the Google Books project's scanning of millions of books—including full-text copying—constituted fair use because it served the novel purpose of searchable excerpts.

Both precedents support the logical chain of "large-scale copying + transformative purpose = fair use," and the DOJ's argument proceeds precisely along this path.

But there are fundamental differences this time. Google Books' outputs were limited to excerpts, and Google established a copyright claim mechanism; LLMs, however, can under certain conditions reproduce training data verbatim—when The New York Times filed its lawsuit, it submitted screenshots of ChatGPT reproducing its articles almost word-for-word as evidence. The DOJ's response is to sever training and output into two independent legal questions, but whether this severance can hold up in court is the core point of contention in the entire case.

Furthermore, the DOJ explicitly criticized the court's analysis in Kadrey v. Meta (2025) in its filing, stating that the court "conflated training with output" when applying the fourth fair use factor, thereby reaching an erroneous conclusion. This criticism carries substantive significance—it not only provides arguments for the current case but also offers directional guidance for the legal interpretation of similar future litigation.

What Is Most Likely to Happen Next

At the court level, Judge Sidney Stein, upon receiving the summary judgment motions on September 4, will face a choice: whether to adopt the government's "training/output dichotomy" framework or to follow the holistic infringement determination path proposed by The New York Times. The government's intervention raises the probability that the court will lean toward the former, but it does not guarantee the outcome.

At the industry level, regardless of how the court rules, this document has already sent a clear policy signal: the Trump administration does not intend to become a regulatory obstacle to AI data compliance. This will accelerate capital flows into AI infrastructure, particularly investments related to data processing and model training, as reduced regulatory uncertainty directly affects risk premiums.

At the global level, the United States' explicit position will intensify policy divergence among jurisdictions. The EU's AI Act imposes stricter requirements on training data transparency, and the divergence in regulatory approaches between China and the United States will become a long-term variable shaping the global layout of the AI industry.