Articles
AI Data8 minute read

Amazon’s Book-Scanning Pipeline Is Also a Preservation Test

Buying a physical copy can resolve one acquisition question. It does not decide whether a scarce artifact should be cut apart, whether the digital surrogate will survive, or whether anyone can audit what entered a model.

An antique book entering a dark industrial scanning line and becoming red data particles

An investigation published August 17 traced a shipment containing rare books to Amazon’s VGT3 facility in Las Vegas. 404 Media placed a tracker in the shipment and reported that workers at the site receive books, remove their bindings, scan the pages, and discard the physical copies. Amazon told the outlet that it purchases books through commercial channels to improve products and services. TechCrunch reported the same statement and described the scans as AI training data.

The reporting establishes more than an anonymous bulk-buying pattern: it identifies a company and a destination through a physical shipment. It does not establish that every book entering VGT3 is rare, that Amazon knowingly targeted a unique surviving copy, or which model or product used a particular scan. Those unanswered questions matter because “rare” can describe market scarcity, historical importance, a specific edition, or simply an out-of-print title.

Ownership, copyright, and preservation are different questions

A company that buys a book owns that physical copy. Copyright still governs reproduction, while preservation asks whether destroying the object erases value that a text-only file cannot capture. Marginalia, bindings, paper, illustrations, printing variations, provenance marks, and evidence of prior ownership can make one copy significant even when another edition contains the same words.

The legal landscape is not reducible to “bought means licensed” or “training means infringement.” In the 2025 Bartz v. Anthropic order, a federal district court treated the one-for-one digitization of lawfully purchased print books, whose source copies were destroyed and digital replacements kept internally, as fair use. The same order treated Anthropic’s retained pirated library differently. The case concerned Anthropic and particular facts; it did not create a preservation standard for every company or settle every AI-training dispute.

The U.S. Copyright Office has likewise analyzed acquisition and training use as fact-dependent questions, including the source of copies, the purpose of use, market effects, licensing, and the nature of outputs. A lawful path to a file can therefore coexist with serious questions about stewardship and disclosure.

A high-quality corpus needs provenance, not just clean text

Printed books published before the current flood of machine-generated material are attractive because they contain edited human writing that may not exist online. But a destructive scan can turn a verifiable artifact into a private digital asset. If the file, metadata, and acquisition record stay inside one company, outsiders cannot check edition accuracy, scanning errors, exclusions, rights status, or whether a supposedly rare source still exists elsewhere.

That is a model-quality problem as well as a cultural one. Optical character recognition can distort tables, footnotes, non-Latin scripts, damaged pages, and unusual typography. Edition metadata affects dates and attribution. Training teams that strip away provenance may later be unable to explain why a model learned an error or reproduce the exact dataset used for an evaluation.

The minimum useful record is item-level: title, author, edition, identifier where one exists, seller, acquisition date, scan method, rights basis, preservation status, file checksum, and every model or dataset version that consumed it. Sensitive commercial details can be access-controlled without making the entire chain unknowable.

Scarcity should trigger a stop rule before the blade

Bulk scanning needs an exception process. Before destructive digitization, operators should check library catalogs, institutional holdings, edition rarity, condition, special markings, and whether a nondestructive scan is practical. A title with abundant equivalent copies may pass quickly. A scarce edition, annotated volume, fragile artifact, or uncertain match should be diverted for expert review rather than treated as interchangeable feedstock.

Companies can publish aggregate sourcing and preservation reports without revealing model recipes. Useful disclosures would include books purchased, copies destructively scanned, copies preserved or donated, rarity checks performed, languages and publication periods represented, correction procedures, and whether qualified libraries can receive digital surrogates where rights permit.

The new investigation shifts the debate from a hypothetical supply chain to a named one. Amazon can answer it by documenting what VGT3 acquires, how scarcity is assessed, what survives after scanning, and who can audit the result. A training corpus should not become more valuable by making its physical sources less available to everyone else.

Quick questions

Did Amazon confirm that it scans books for AI training?

Amazon said it purchases books through commercial channels to improve products and services. 404 Media traced a shipment to VGT3 and reported, based on workers and its investigation, that the facility destructively scans books for AI training.

Is scanning a purchased book for AI training automatically legal?

No single rule resolves every case. A federal district court found specific uses of lawfully purchased and destructively scanned books fair in Bartz v. Anthropic, but acquisition, copying, retention, outputs, and market effects remain fact-dependent.

Why does destructive scanning matter if a digital copy remains?

A file may preserve text while losing physical evidence such as edition details, marginalia, paper, bindings, and provenance. Private files can also disappear or remain unavailable to libraries and researchers.