Reporters at 404 Media tracked a shipment of rare and used books through the resale supply chain and found it ended up at an Amazon facility, where the books were scanned and then destroyed, according to the outlet’s investigation published this week.

How the Investigation Worked
To find out what happens to books after they are sold in bulk to online resellers, 404 Media purchased a batch of used and rare titles, hid tracking devices inside the packaging, and monitored the shipment’s journey. The trail led to a facility operated by Amazon, where the outlet reported that books are processed through scanning equipment before being discarded or destroyed.
The report does not claim direct proof that the scanned text from these specific books was fed into an AI model, but it places renewed scrutiny on Amazon’s book-handling operations at a moment when every major tech company is racing to acquire more high-quality text data to train large language models. Publishers, authors, and researchers have said for years that books represent some of the richest training material available because of their length, coherence, and edited quality compared to scraped web content.
Why Book Data Matters for AI
Large language models are trained on enormous volumes of text, and books are considered especially valuable because they contain long-form, well-structured writing that is harder to find at scale on the open web. That value has already triggered a wave of litigation. Authors and publishers have sued OpenAI, Meta, and other AI developers over allegations that copyrighted books were used to train models without permission or compensation, often sourced from pirated datasets such as the one known as Books3.
404 Media’s report suggests a different, less discussed pathway: physical books purchased through ordinary resale channels, digitized at scale, and then discarded. If accurate, it would mean AI developers are not only relying on scraped internet text or leaked datasets, but potentially on physical book stock funneled through commercial supply chains and converted into digital training material.
The investigation underscores how difficult it has become to trace where AI training data actually originates, even when the physical source material can be tracked door to door.
Amazon’s Role in the Book Supply Chain
Amazon has long operated logistics and processing facilities that handle used books sold through its marketplace and trade-in programs, as well as through its ownership of book-focused platforms. The company has also invested heavily in its own AI ambitions, including large language models developed for its Alexa assistant, AWS cloud AI services, and internal research efforts competing with OpenAI, Google, and Anthropic.
Amazon has not published detailed public disclosures about the specific sources of text data used to train its AI systems, a common industry practice that has drawn criticism from researchers and rights holders who argue companies should be more transparent about the provenance of training material. 404 Media said it reached out to Amazon for comment on the shipment and the fate of the scanned books as part of its reporting.
Part of a Wider Pattern
The findings arrive amid growing public unease over how tech platforms acquire the data that powers AI products, often from users or creators who never explicitly agreed to it. Streaming platform Twitch faced backlash earlier this year after streamers discovered contract language that could allow their content to be used for AI training, prompting the company to respond publicly, as NarwhalTV reported. That controversy, like the Amazon book findings, highlighted how AI training practices are frequently uncovered only after independent scrutiny rather than proactive disclosure.
The broader AI industry has also faced pressure over how it balances rapid growth with public trust. Anthropic CEO Dario Amodei has argued that AI companies need to deliver undeniable public benefits, such as breakthroughs in disease treatment, to justify the technology’s disruptive costs, a point covered in depth by NarwhalTV. Episodes like the alleged destruction of scanned books after AI processing complicate that trust-building effort, particularly for authors and small booksellers who may have no visibility into where their physical inventory ultimately lands.
What Happens Next
404 Media’s report does not indicate that any laws were broken, since buying used books in bulk and scanning them is not inherently illegal, and copyright law around AI training remains an unsettled legal question working its way through multiple federal courts. But the investigation is likely to add fuel to ongoing lawsuits and congressional interest in AI transparency requirements.
For authors and publishers, the report reinforces long-standing concerns that once a physical book enters a resale or liquidation pipeline, its content may be digitized and repurposed with little recourse for the original rights holder. Advocacy groups such as the Authors Guild have pushed for clearer disclosure rules and licensing frameworks that would compensate writers when their work is used to train commercial AI systems.
Amazon has not issued a detailed public statement addressing the specific claims in the 404 Media report as of publication. NarwhalTV will continue to follow developments as more details emerge about the scale of the practice and whether it extends beyond the single shipment tracked by reporters.