A bookseller investigating a mysterious bulk buyer slipped an AirTag into one of roughly 1,000 books. The tracker led to Amazon facilities in Las Vegas, where workers told 404 Media that books are scanned, stripped of their bindings and fed page by page into document scanners.
The process creates a digital copy by destroying the physical one.
Amazon did not dispute that it acquires books for product development. “Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” a spokesperson told 404 Media. The company did not say which products receive the resulting text.

The investigation is the first reported evidence connecting Amazon to the large, unexplained book orders that sellers had been receiving. Its significance is not that one strange shipment disappeared into a warehouse. It is that the shipment reveals an industrial method for obtaining human-written material that may not exist online.
A physical library becomes machine-readable data
Workers described receiving the books, scanning their ISBN barcodes, removing their bindings and scanning the loose pages. The route included Amazon’s LAS8 warehouse and its VGT3 operation, whose team logo depicts a dinosaur biting into a book.
The books have widely been described as “rare,” but that word needs care. Here it means that few copies remain in circulation, not necessarily that the shipment contained valuable first editions. Because 404 Media withheld the titles to protect its source, outsiders cannot judge what was lost or how easily those copies could be replaced.
Earlier suspicious orders included obscure nonfiction from the 1970s through the 1990s, as well as specialized academic works. These are precisely the kinds of texts that may be commercially unremarkable yet useful to a developer seeking material beyond the modern web.
The buy-one, destroy-one model has a legal precedent
Amazon is not the first AI company linked to destructive book scanning. In Bartz v. Anthropic, a federal court found that Anthropic spent millions of dollars buying millions of print books, then paid vendors to remove their bindings, scan their pages and discard the originals.
The judge held that converting each lawfully purchased book into one internal digital copy was fair use. The logic was unusually concrete: one purchased physical copy became one digital copy, without increasing the number of copies in circulation.
“The print original was destroyed. One replaced the other.”
That ruling did not give AI companies a blanket exemption from copyright law. The court separately refused to excuse Anthropic’s acquisition of pirated books. But it established a plausible legal route for the method Amazon now appears to be using: buy a copy, digitize it and destroy the original.
Why old books matter now
AI developers need enormous quantities of genuine human writing. That material becomes more valuable as the public internet fills with text produced by AI systems themselves.
A peer-reviewed Nature study found that repeatedly training models on generated output causes errors to accumulate and uncommon information to disappear. In practical terms, models fed too heavily on their own descendants can gradually lose contact with the full range of the original human record.
Books that never made it onto the web therefore represent more than old inventory. They are reservoirs of human-origin language and knowledge. Amazon markets its own Nova family of foundation models and tools for customizing models, but it has not identified whether those systems use the scanned books.
The contrast with Amazon’s origins is sharp. In 2003, the company introduced Search Inside the Book using 120,000 titles supplied through participating publishers. Digitization was presented as a way to help people discover and purchase physical books.
The operation uncovered in Las Vegas reverses that relationship. The book is no longer the product being preserved and sold. It is the raw material consumed to produce something else.