Key takeaways
- An AirTag placed in a bulk book order confirmed Amazon's Las Vegas facility is physically destroying rare books to scan them.
- Workers scan ISBNs to systematically ingest un-digitized works, securing proprietary data absent from web scrapes.
- The practice exposes the severe training data crunch leading frontier AI labs to exploit niche, offline text caches.
What happened
An investigative report by 404 Media exposed that Amazon operates a dedicated facility in Las Vegas focused on bulk-buying rare and out-of-print physical books, ripping them from their bindings, and scanning the pages to create training data for its foundation models. The operation was uncovered after a bookseller placed an Apple AirTag inside a bulk shipment, which pinged from Amazon’s VGT3 warehouse. Internal facility logos reportedly featured a dinosaur devouring a book, underscoring the destructive nature of the throughput.
Discussions from warehouse workers revealed that staff are instructed to log barcodes and ISBNs prior to slicing the spines for high-speed page scanning. This methodical indexing supports longstanding theories across the antiquarian book trade that AI developers are working through historical ISBN registries to absorb every available physical volume. Workers also noted intermittent supply shortages where operations nearly paused due to a lack of unscanned physical literature.
In response, Amazon issued a general statement stating it purchases books through commercial channels to develop and enhance customer products and services, declining to specifically address AI training. While competitors like Anthropic and xAI have claimed they do not ingest antique books, the findings show Amazon actively seeking proprietary offline corpora to keep pace with frontier model developers like OpenAI and Google.
Why it matters
This disclosure underscores the extreme measures frontier AI developers are taking to overcome the looming data wall. As public internet scrapes become saturated, heavily synthesized, or legally contested, un-digitized physical books represent one of the last high-density reserves of clean, human-written natural language.
By physically destroying the media after scanning, Amazon secures proprietary access to unique linguistic patterns and niche knowledge that cannot be replicated by rivals relying strictly on web-crawled datasets.
Beyond data strategy, the practice introduces significant reputational and ethical friction for corporate AI programs. Destroying rare, historical, or foreign-language works solely for private model weights sparks cultural backlash from preservationists and content creators. Furthermore, because these scanned texts remain private to avoid leaking training advantages, valuable cultural and historical knowledge is permanently locked behind corporate model checkpoints rather than shared through open digital archives.
What to watch
Keep an eye on how regulatory bodies and copyright frameworks respond to physical-to-digital training pipelines where physical copies are legally purchased but destructively ingested for commercial LLMs. Additionally, watch whether booksellers organize collective restrictions against bulk AI brokers, and monitor whether rival frontier labs expand similar offline data collection pipelines or face heightened pressure to disclose the origins of their training corpuses.




