US Department of Justice Takes First Public Stance on AI Training: Trump Administration Backs Fair Use in New York Times v. OpenAI Copyright Case

On September 2, 2026, the US Department of Justice filed a statement of interest in the New York Times v. OpenAI copyright case, formally declaring that training large language models on copyrighted text constitutes fair use. This marks the first time the federal government has taken a formal position in AI-related copyright litigation.

On September 2, 2026, the US Department of Justice filed a 20-page "statement of interest" with the federal court in Manhattan, formally establishing the federal government's legal position: training large language models on copyrighted text constitutes fair use and does not violate copyright law. Co-signed by Solicitor General Stanley Woodward Jr., the filing was submitted to presiding Judge Sidney Stein in the US District Court for the Southern District of New York, where the case is being heard. This marks the first time the US federal government has formally weighed in on the series of copyright lawsuits involving AI companies.

The document intervenes in the copyright suit that The New York Times filed against OpenAI and Microsoft in December 2023. The Times alleges that the two companies used millions of published articles without permission to train the large language models underlying ChatGPT, and that generated content competes directly with its news products, eroding subscription, licensing, and advertising revenue. OpenAI denies infringement, arguing that training on publicly available materials is protected by the fair use doctrine under long-standing case law and is essential to maintaining US AI competitiveness. Lawsuits filed by several affiliated newspapers have been consolidated into this lead case; in June of this year, nearly 400 news publishers filed a separate class action against OpenAI and Microsoft making the same allegations.

The technical fair-use dispute: what LLMs actually do with copyrighted content

Fair use is an exception embedded in US copyright law that permits otherwise unauthorized use of protected works under certain circumstances. Four factors guide the inquiry: whether the purpose of the use is commercial and "transformative" in character; the nature of the copyrighted work; the amount used relative to the work as a whole; and the effect of the use on the market for the original. The factors must be weighed together — none alone is dispositive, and no fixed formula exists. The outcome rests entirely on case-by-case judicial discretion.

The Department's core technical argument is that the process of training a large language model is fundamentally different from ordinary copying. The model does not directly store or reproduce text — training data is transformed into hundreds of millions of floating-point parameters that encode statistical regularities of language, not the substance of any particular article. At inference time, every passage the model generates is a probabilistic sample conditioned on those parameters; the specific wording of an original article cannot, in theory, be reconstructed from the model. The filing describes this process as "extraordinarily transformative," asserting that what the system derives from training data is "general reasoning and language ability," not replayable original content.

The Times's rebuttal targets precisely the weak point in that account. The paper has presented evidence in the litigation showing that, under certain prompts, ChatGPT can reproduce passages from its paywalled articles nearly verbatim — demonstrating that the model can, under some conditions, memorize specific content rather than merely absorb abstract linguistic patterns. This "memorization spillover" is a recognized practical problem in AI engineering. It implicates not only the training process itself but also output controls at the deployment stage: even if training is lawful, reproducing content at generation time could constitute independent infringement.

How much influence the document carries, and why the timing matters

Reading the legal force of a statement of interest accurately is the prerequisite for measuring the true significance of the filing. These documents are a standard mechanism by which federal agencies express a position without intervening in a case: the government does not become a party, the court is not obligated to adopt its view, and full decisional authority remains with Judge Stein. Yet the document carries weight. A formal submission co-signed by the Solicitor General and invoking national security and economic competitiveness provides an interpretive framework the court must account for when weighing "public interest" as a fair-use factor — and on appeal, a formal federal government position would receive even more careful consideration from higher courts.

The choice of timing appears calculated. Judge Stein had ordered both sides to file summary judgment motions by September 4. The Department submitted its filing two days earlier, injecting a policy signal directly into the most consequential phase of the litigation.

The government's argument structure: from innovation narrative to national security

The statement of interest rests not on copyright doctrine alone but on a framework that converts the copyright question into a matter of national strategy. Its argument proceeds on three levels: that large language models are yielding major breakthroughs across research fields; that a cramped reading of fair use would impede creative and scientific progress; and that if copyright rules substantially raise the difficulty of training large models in the United States, the competitiveness of the American AI industry would suffer, creating national security risks. The filing states the position without ambiguity: "The United States has a strong interest in this case in requiring the Court to reject any claim that training LLMs on copyrighted text violates copyright law."

The maneuver lifts the copyright issue out of the intellectual-property framework and repositions it within geopolitical competition — a direct echo of the core narrative Silicon Valley AI companies have long pressed in Washington: that with US-China AI competition intensifying, the cost of copyright restriction is not merely commercial but strategic. The Times responded forcefully. "The government is siding with a few trillion-dollar AI companies at the expense of the countless American creators whose work those companies have stolen," a spokesperson for the paper said, adding: "The administration's proposal would undermine the sustainability of the human-authored content that a healthy society depends on — and on which AI itself depends to function."

What each set of stakeholders stands to gain or lose

For AI developers — OpenAI, Meta, Anthropic and others — the government's entry provides a meaningful layer of political risk protection, even absent direct legal binding force. Copyright exposure is the central uncertainty in these companies' business models: a judicial finding that training infringes could, in principle, reach back across all historical training data and produce enormous damages. The federal government's open endorsement of the fair use position at minimum lowers the expected probability of that scenario. Microsoft, OpenAI's largest financial backer, benefits equally, as its legal exposure overlaps almost entirely with OpenAI's.

For the news publishing industry, the filing marks an unmistakable political setback. Publishers had hoped the government would remain neutral, if not sympathetic, on copyright. The White House's ordering of priorities is now explicit: AI industrial competitiveness precedes the protection of content rights. The class action of nearly 400 news publishers and the petition signed by more than 400 Hollywood creators and executives — both directed at the current administration — are implicitly set aside by this position.

For book publishers, music rights holders and film and television creators, the consequences extend far beyond news media. AI training corpora span books, lyrics, scripts and visual art, and the Department's framing, built around "publicly available internet materials," is general in its reach rather than confined to newspaper articles. Hollywood's unions and studios have already conducted protracted labor negotiations over AI's use of copyrighted works; this political signal will directly reshape their assessment of bargaining leverage.

For enterprise developers and SaaS users, near-term policy uncertainty has declined.

The divergent 2025 rulings as legal reference points

In 2025, two federal judges reached conflicting conclusions on analogous fair-use questions, and no appellate court has yet supplied a uniform standard. In a case involving Anthropic, a California federal judge held that Anthropic's use of authors' works to train AI models was "extraordinarily transformative" and thus fair use; the same court separately held that Anthropic's use of pirated materials to build a digital content library did not qualify as fair use and ordered that claim to trial — the case ultimately settled. Another judge, weighing market harm, reached a different conclusion. The root disagreement concerns how much weight harm to the market for original works should carry once a use has been found transformative.

This precedent landscape matters directly for the Times case. The Times's complaint involves no pirated materials; its claim is the more fundamental one that the large-scale copying of lawfully accessed public web content for training is itself infringing. That is the substantive target of the Department's intervention. With no controlling circuit precedent on the question, Judge Stein's ruling would become the first benchmark for the industry, cited widely across dozens of pending and future suits.

Strategic outlook

The most likely near-term development is the pivotal decision at the summary judgment stage. Both sides filed their motions on September 4; Judge Stein's ruling will be the first decisive milestone. If Stein adopts the "extraordinarily transformative" framework, copyright risk across the AI industry narrows dramatically and a developer-friendly industry standard takes practical hold. If he instead focuses on the evidence of verbatim output, he may require OpenAI to prove that its output control mechanisms sufficiently suppress content reproduction — a requirement that would force every AI company to accelerate engineering investment in unlearning techniques and output filtering, with the costs transmitted through the entire development ecosystem.

Further out, the decisive signal will come at the circuit level. The trial ruling in this case will eventually reach the Second Circuit, which has jurisdiction over the Times suit, while parallel cases are advancing in the Ninth Circuit as well. If the two circuits diverge, pressure for Supreme Court review will intensify sharply. That would be the true final adjudication, fixing the boundaries of fair use in the AI training context once and for all. For now, the federal government has made clear, across 20 pages, where it wants those boundaries drawn.

On the legislative front, the filing transmits a clear message: the current administration has no intention of narrowing fair use through statute. That relieves lobbying pressure on AI companies while leaving copyright holders with essentially no avenue other than litigation. The practical consequence is that copyright suits will not diminish, and highly uncertain outcomes will form the ambient condition of the content industry for years to come.