Anthropic's "Panama Project" involved purchasing millions of used physical books, which were scanned and then destroyed to train the Claude model. Court records show that the project was described in internal planning documents in 2024 as requiring confidentiality, with the documents stating, "We don't want the outside world to know we are doing this."
Fact Restoration
According to unsealed documents reported by The Washington Post, Anthropic initially downloaded digital books from LibGen in 2021 before later shifting to physical books. The project used hydraulic cutters to remove book spines, high-speed scanners to digitize pages, and ultimately had recycling companies process the remaining books. In August 2025, the company reached a $150 million settlement with a class-action lawsuit from authors, with the judge ruling that the use constituted transformative use and fell under fair use.
These operations relied on the first sale doctrine, which allows buyers to dispose of purchased physical books without interference from copyright holders. Documents show that Anthropic viewed physical books as a "necessary" data source to avoid the model outputting "low-quality internet language."
Mechanism Breakdown
From a business logic perspective, the cost of acquiring physical books was lower than licensing fees, and physical destruction reduced the risk of claims regarding additional copies. Technically, cutting and high-speed scanning enabled large-scale processing, with internal documents mentioning plans to expand to "all books in the world." This approach met training needs while attempting to operate in a legal gray area.
The company initially relied on pirated digital libraries but later shifted to legally purchasing physical books, reflecting a gradual adjustment in data source compliance. Unsealed documents reveal that leadership viewed book content as key to improving model writing quality.
Industry Impact
For competitors, this event may prompt other AI companies to reassess data acquisition strategies to avoid similar legal risks. Upstream and downstream publishers and author groups received settlement compensation, but future licensing negotiations may become more cautious.
Developers and enterprise users face the possibility of rising costs for training data acquisition. If similar practices are restricted, the iteration speed of models relying on public or licensed data may be affected.
Comparison and Precedent
Compared to earlier AI companies scraping data from the web, Anthropic's shift to physical book destruction reflects a strategic change as scale expanded. Historically, the first sale doctrine supported the used book market; this case extends it to AI training, with the judge ruling that no additional physical copies were created or redistributed.
Strategic Assessment
Based on the existing documents, this case may lead more AI companies to adjust their data strategies. The incident highlights the conflict between AI training's reliance on high-quality text and copyright boundaries, with future data acquisition likely shifting more toward licensed channels.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接