Amazon Is Reportedly Buying and Destroying Books to Feed Its AI

Books are scanned into digital data for AI training.

Books are scanned into digital data for AI training. Image: Generated via Google’s Nano Banana

Written By
Matt Gonzales
Matt Gonzales
Aug 17, 2026
5 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

A rare book went into the mail. What came out the other end was a glimpse inside the increasingly physical race for AI training data.

Amazon is buying large quantities of printed books, scanning them for AI training data, and destroying the physical copies in the process, according to a 404 Media investigation. The publication uncovered a previously unreported operation: hiding a tracking device inside a rare book suspected of having been purchased for AI training, then following the device to an Amazon facility in Las Vegas.

The finding adds Amazon to an industrywide hunt for high-quality human-written material as AI developers look beyond the open web for data to train increasingly sophisticated models.

A tracked book led reporters to Amazon

404 Media's investigation began with a decidedly low-tech experiment.

Reporters placed a tracker inside one of several rare books shipped by a bookseller that had experienced a surge in unusual orders. The device eventually led them to VGT3, an Amazon facility in Las Vegas.

According to Amazon employees interviewed by 404 Media, workers at the facility receive large shipments of printed books and remove their bindings so the pages can be scanned more quickly. The process destroys the physical books.

Amazon confirmed that it is acquiring books for its technology efforts.

“Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” a company spokesperson told 404 Media.

The company did not tell the publication which specific AI products are trained using the scanned material.

That leaves an important wrinkle in the story. The investigation does not accuse Amazon of obtaining the tracked books from pirate libraries or illegally downloading them. Instead, it documents a physical pipeline in which books are commercially purchased, scanned, and destroyed as their contents are converted into digital data.

Advertisement

Why AI companies want old books

Amazon's operation arrives amid a broader push by AI developers to find large quantities of high-quality human-created text.

In an earlier investigation, 404 Media reported that AI companies were buying old printed books in part because books published before the generative AI boom offer material that predates the rapid growth of AI-generated content online. That distinction has become increasingly valuable as developers consider the quality and provenance of their training data.

Books also offer something the sprawling internet cannot always guarantee: edited, structured writing across specialized subjects, languages, professions, and historical periods.

For AI companies hungry for more training material, the world's bookshelves become another potential data source. But converting those shelves into datasets can come at a physical cost.

The books tracked by 404 Media were described as rare because relatively few copies remained in circulation. The publication did not reveal the titles, noting that scarcity could stem from limited print runs or from books being written in languages with relatively few speakers.

That does not necessarily make every book a valuable collectible. It does mean that destroying a copy can permanently reduce an already limited supply.

The Amazon revelation lands amid an escalating fight over whether and how copyrighted books can be used to develop generative AI.

In July, Hachette Book Group, Cengage Learning, Elsevier, and author Scott Turow filed a proposed class-action lawsuit accusing Google of improperly using copyrighted books and journal articles to develop Gemini.

As eWeek reported on the Gemini copyright lawsuit, the case raises a particularly thorny question: If publishers provided material to a technology company for services such as search or ebook sales, does that permission extend to training an AI model? The publishers argue that it does not.

Anthropic has faced a separate and closely watched legal fight over its use of books.

In a 2025 ruling, US District Judge William Alsup concluded that Anthropic's use of copyrighted books to train its AI models qualified as fair use. But the judge treated the acquisition of the books as a separate issue, ruling that Anthropic's creation of a central library containing pirated copies was not protected simply because the material could later be used for AI training.

Advertisement

The litigation eventually produced a massive settlement over the pirated works. In July, a federal judge gave final approval to Anthropic's $1.5 billion copyright settlement, covering roughly 500,000 works and allocating about $3,000 per title before fees and other costs.

The Amazon investigation does not establish that the company is following the same path that triggered the Anthropic litigation. Amazon told 404 Media that it purchases its books through commercial channels.

Instead, the operation illustrates another unresolved piece of the AI copyright puzzle: legally acquiring a physical copy of a book does not, by itself, answer every question about what rights accompany its purchase or how its contents may subsequently be used.

The larger legal landscape remains unsettled. The US Copyright Office has resisted treating all AI training as automatically fair or infringing, instead emphasizing that traditional fair-use analysis depends on the circumstances surrounding how copyrighted material is acquired and used.

As eWeek's coverage of the Copyright Office's AI training findings explains, questions surrounding licensing, fair use, and the economic effect of AI training remain far from resolved.

When digitization means destruction

The Amazon investigation also raises a question that goes beyond copyright: What happens to the books themselves?

Digitization has historically been closely associated with preservation. Libraries and archives scan fragile material so information can survive even if the original eventually deteriorates. Destructive scanning reverses that equation. Removing a book's binding allows pages to move through high-speed scanners efficiently, but the process sacrifices the physical copy.

For widely available books, the loss of an individual copy may have little effect on overall availability. The equation becomes more complicated when the same process reaches books with limited print runs, obscure foreign-language works, out-of-print technical references, or other titles with relatively few surviving copies.

404 Media's investigation does not establish how many scarce books Amazon has destroyed or identify the titles it has processed. Those unanswered questions matter when considering the potential impact on the physical book supply. What the investigation does establish is that Amazon has built an operation capable of converting commercially purchased books into digital format while destroying the originals. And that exposes a strange new economics of knowledge.

Advertisement

For years, AI developers could treat the internet as an enormous reservoir of training material. Now, as AI-generated text occupies a growing share of the digital world, reporting suggests that human-written material from before the generative AI boom has acquired a new kind of value.

Also read: As Amazon looks for new sources of high-quality training data, the company is also overhauling its Nova AI strategy, signaling a broader push to strengthen its position in the increasingly competitive AI race.

Matt Gonzales

Matt Gonzales is the Managing Editor of Cybersecurity for eSecurity Planet. An award-winning journalist and editor, Matt brings over a decade of expertise across diverse fields, including technology, cybersecurity, and military acquisition. He combines his editorial experience with a keen eye for industry trends, ensuring readers stay informed about the latest developments in cybersecurity.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.