According to an investigation by 404 Media, Amazon has been purchasing rare books in bulk, removing their spines, and scanning the pages for use in AI training. 404 Media placed a tracking device inside a rare book that ultimately arrived at an Amazon facility in Las Vegas identified as VGT3.
The VGT3 site is said to be marked with an emblem of a dinosaur holding a book. In a statement to 404 Media, Amazon said it “purchases books through commercial channels to improve the products and services customers use.”
Why these books matter for AI training
Companies that train large language models (LLMs) require enormous quantities of text. While much material has already been harvested from the open web, out-of-print and otherwise hard-to-find books provide additional, valuable training material because they are not available online.
An important point cited by the report is that anything published before 2022 could not have been written by an LLM, so such texts help avoid contaminating training data with AI-generated content. Training on large amounts of AI-produced text can risk a phenomenon known as “model collapse,” where a model’s output quality degrades after ingesting too much machine-generated text.
Related concerns
The report also references broader controversies about how companies obtain book content. It notes prior claims that other firms, for example Anthropic, used books acquired through improper means. The practice of unbinding and scanning physical books raises questions about provenance, copyright compliance and the composition of training datasets.
What to watch next
The story highlights issues of traceability, legal and ethical use of copyrighted works, and how big tech sources training data. Further investigations, corporate disclosures and potential legal scrutiny are likely to provide more detail on the scope and legality of this acquisition and digitization activity.



