Anthropic Destroys Rare Books to Train Claude Models, Cultural Destruction Controversy Intensifies

Leaked tweets from July 29, 2026 reveal that Anthropic purchased millions of physical books from bookstores and wholesalers, scanned them via optical character recognition, then destroyed all copies, retaining only standardized training text fragments. The process sparked widespread backlash over data ethics and cultural preservation.

Leaked tweets from July 29, 2026 reveal that Anthropic purchased millions of physical books—covering fiction, academic works, and professional textbooks—from bookstores and wholesalers via third-party suppliers. After extracting text using optical character recognition technology, the company uniformly destroyed all physical copies, retaining only the standardized training text fragments.

This process formed a complete chain of procurement, digital extraction, and destruction of physical media. Internal documents indicate that publicly available web text data suffers from inconsistent quality and high copyright risks, whereas book content offers higher authority and better structure, improving model reasoning ability and knowledge accuracy. The project reportedly began as early as 2022 and continued through the end of 2023, primarily conducted within the United States, with scanning handled by a partner technology service provider.

Operational Mechanism and Business Logic

Anthropic, founded in 2021 by former OpenAI employees, was valued at approximately $15 billion in 2023. Its core product is the Claude series of large models. By destroying physical books, the company sought to downplay the copyright implications of copying behavior, avoiding the high costs and complex processes of obtaining direct authorization. In contrast, OpenAI entered into copyright partnerships with publishers such as Penguin Random House and HarperCollins in 2023, securing legal authorization from the source. Google DeepMind released a training data transparency white paper in March 2024, detailing data sources and licensing status.

Companies like ISBNdb offer bulk book services, claiming to hold the world’s largest book database with 111 million records and providing catalogs of rare professional volumes. Bookseller feedback indicates that weekly sales jumped from around 20 copies to hundreds, primarily for hard-to-find out-of-print books.

Industry Impact Analysis

Regarding the competitive landscape, this incident highlights differences in data acquisition methods. Anthropic’s approach sparked backlash on X platform. Michael Burry called it “the embodiment of evil,” Elon Musk demanded that SpaceX’s AI team scan books non-destructively, and David Sacks pointed out double standards in training for free and IP protection. In contrast, OpenAI and Meta have pursued compliance-oriented paths. Meta launched the Open Data Alliance in June 2024, partnering with publishers and academic institutions to build a legal shared training data platform.

In terms of upstream and downstream impact, European and American booksellers have received orders from Singapore and elsewhere for 3,000 rare English-language books. Some antique book dealers worry that uncommon books are being turned into pulp. The EU AI Act, effective May 2024, requires developers to disclose training data sources and copyright status, with penalties of up to 4% of global annual turnover for violations. McKinsey’s 2024 AI Data Ethics Report shows that 68% of AI companies have opaque data sources, and 32% face potential copyright risks.

For developers and enterprise users, compliance paths are costlier but reduce litigation risks. Anthropic has agreed to pay $1.5 billion to settle a class-action lawsuit involving hundreds of thousands of pirated books—the largest copyright settlement in U.S. history. The judge ruled that transformative use is protected by fair use, but the case did not resolve all disputes.

Strategic Assessment

Based on available facts, the most likely next development is an increase in lawsuits targeting similar data destruction practices.