Safety

AI-generated text

Amazon reportedly buys rare books, cuts spines and scans them for AI training

Investigative reporting by 404 Media found that Amazon has purchased rare books which were sent to a Las Vegas facility—identified as VGT3—where books were reportedly unbound and scanned for use in AI training.

Amazon reportedly buys rare books, cuts spines and scans them for AI training

According to an investigation by 404 Media, Amazon has been purchasing rare books in bulk, removing their spines, and scanning the pages for use in AI training. 404 Media placed a tracking device inside a rare book that ultimately arrived at an Amazon facility in Las Vegas identified as VGT3.

The VGT3 site is said to be marked with an emblem of a dinosaur holding a book. In a statement to 404 Media, Amazon said it “purchases books through commercial channels to improve the products and services customers use.”

Why these books matter for AI training

Companies that train large language models (LLMs) require enormous quantities of text. While much material has already been harvested from the open web, out-of-print and otherwise hard-to-find books provide additional, valuable training material because they are not available online.

An important point cited by the report is that anything published before 2022 could not have been written by an LLM, so such texts help avoid contaminating training data with AI-generated content. Training on large amounts of AI-produced text can risk a phenomenon known as “model collapse,” where a model’s output quality degrades after ingesting too much machine-generated text.

Related concerns

The report also references broader controversies about how companies obtain book content. It notes prior claims that other firms, for example Anthropic, used books acquired through improper means. The practice of unbinding and scanning physical books raises questions about provenance, copyright compliance and the composition of training datasets.

What to watch next

The story highlights issues of traceability, legal and ethical use of copyrighted works, and how big tech sources training data. Further investigations, corporate disclosures and potential legal scrutiny are likely to provide more detail on the scope and legality of this acquisition and digitization activity.