<p id="speakable-summary" class="wp-block-paragraph">Amazon is buying tons of rare books, cutting off their spines, and scanning them for AI training, according to <a rel="nofollow" href="https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/">404 Media</a>, which placed a tracking device in a rare book that ultimately arrived at an Amazon facility in Las Vegas.</p>
<p class="wp-block-paragraph">The facility, known as VGT3, identifies itself with a symbol of a dinosaur holding a book in its claws. Amazon told 404 Media in a statement that it “purchases books through commercial channels to improve the products and services customers use.”</p>
<p class="wp-block-paragraph">Companies like Amazon need unfathomably large amounts of text to train their LLMs, which have already ingested what they can from the internet (and, in Anthropic’s case, illegally <a href="https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved/">pirated</a> books). Rare books, especially ones that are out of print or impossible to find on the internet, offer a new source of coveted training data.</p>
<p class="wp-block-paragraph">These texts are especially valuable since there’s no chance that anything published before 2022 was written by an LLM. When LLMs train on AI-generated text, they risk “<a href="https://techcrunch.com/2024/07/24/model-collapse-scientists-warn-against-letting-ai-eat-its-own-tail/">model collapse</a>,” which can occur when the quality of an LLM’s outputs degrade after ingesting too much AI-generated text.</p>
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接