Amazon confirmed it purchases books through commercial channels to develop and improve its products and services, following a report that traced a bulk order of rare books to a facility where workers reportedly destroy and scan physical copies for digitization.

The practice highlights the intense competition and evolving methods tech companies are employing to acquire vast amounts of data for training increasingly sophisticated AI models, raising questions about copyright, fair use, and the future of digitized literature.
Amazon has stated that it purchases books through commercial channels to develop and improve its products and services, following an investigation that tracked a bulk order of rare books to an Amazon facility where workers reportedly destroy and scan physical copies.
The investigation by 404 Media tracked an order of approximately 1,000 books to an Amazon facility in Las Vegas using a hidden Apple Airtag. Workers at the site, known as VGT3, reportedly remove the spines of books before scanning their pages, rendering the physical copies unusable. Employee posts also indicated that workers scan the books’ barcodes or ISBN numbers as part of the process.
An Amazon spokesperson told City AM that the company buys books to "help develop and improve the products and services our customers use." However, Amazon did not specify which products the material was being used to develop or explicitly confirm that the scanned books were being used to train its AI models.
This practice comes amid a growing demand for high-quality written material among tech companies developing AI systems. Printed and out-of-print books are particularly valuable as much of their content is not available online and predates the proliferation of AI-generated text, which complicates sourcing reliable human-written training data. Amazon is developing its own family of Nova AI models and is competing with companies like OpenAI, Google, and Anthropic in the AI race.
Anthropic, the developer of Claude, previously engaged in a similar practice under its internal program, Project Panama, where it bought millions of second-hand books, cut off their spines, scanned the pages, and recycled the originals. A US federal judge subsequently ruled that Anthropic’s digitization of legally purchased books constituted fair use, distinguishing it from the company’s separate use of pirated copies. Anthropic later agreed to a $1.5 billion settlement covering claims related to pirated books.
Reports from rare book dealers indicate unusually large inquiries from anonymous buyers seeking specialist titles, raising concerns that AI developers or their suppliers are systematically targeting obscure and out-of-print books that have not been digitized. The demand for fresh data is increasing as the easily accessible internet has already been extensively scraped. Publishers are beginning to monetize this demand, with Amazon reportedly striking an AI licensing agreement with The New York Times worth between $20 million and $25 million annually.