亚马逊批量购书扫描用于 AI 训练后销毁
Core Highlights
The tech publication 404 Media ran a tracking experiment that, for the first time, confirmed an undisclosed book-buying operation run by Amazon: the company bulk-purchases large numbers of books, scans them into material for AI training, and then destroys the physical copies afterward. The report turns a rumor that had long lingered in hearsay — that books were being fed to models — into a fact backed by location data rather than by speculation or by someone's anonymous claim posted to a forum. For people who care about content rights, the piece lands like a warning bell that the supply chain for training data is wider and stranger than assumed, and that the books on a shelf may be closer to a data center than to a reader.
What It Does or What Happened
The journalists' method was strikingly hands-on. They hid a tracking device inside a rare book and then let that book drift into the market through ordinary channels, exactly as a normal secondhand sale would. The logistics trail showed the book first entered Amazon's acquisition system and was later delivered to one of the company's artificial-intelligence training facilities. The entire chain of custody is laid out with location data, so this is documentation rather than guesswork: not an anonymous leak, but the literal path a device actually walked from a bookstore shelf to a data center loading dock, with timestamps that line up and remove any doubt about the destination of the shipment in question.
Technical Details
The crucial detail the report surfaces is "scan then destroy." After a book is digitized, it does not flow back into the secondhand market; it is disposed of. That means Amazon's route to acquiring training corpus is the industrial-scale pattern of "buy, scan, discard," not merely a reliance on already-public digital copies that anyone could license or crawl. The volume involved, and the footprint it leaves on the physical book market, are far larger than outsiders had imagined when they assumed training data came mostly from the open web, and the operation leaves almost no public trace to investigate. This is precisely the kind of pipeline that is nearly impossible to audit from the outside, because the evidence is destroyed as soon as it is used.
Versus Competitors
Compared with players like OpenAI and Meta, who lean more on web crawling and licensed libraries, Amazon's path of "take in physical books, then destroy them" is far more concealed and far easier to keep out of the public eye, because the destroyed physical object leaves no secondhand circulation trail that a researcher could trace back to the buyer and expose. Where a crawled webpage is visible to anyone, a shredded book simply disappears, which is precisely what makes the practice hard to detect from the outside and hard to challenge after the fact, even once the public knows it is happening.
Industry Impact or Use Cases
This episode pushes the gray area of AI training-data sourcing one step further into the light. For publishers, authors, and secondhand booksellers, the question of how their content is quietly absorbed into large models, and whether it should be compensated, is becoming an unavoidable issue of copyright and ethics that the industry has so far left largely unanswered, and it may well attract new regulatory attention from lawmakers already uneasy about opaque data pipelines that no one outside the company can inspect or audit. The lesson for the rest of the field is uncomfortable: the most valuable training data may be the kind nobody can prove was ever taken. The uncomfortable implication is that provenance tracking, not just licensing, now sits at the center of AI data ethics. If a book can vanish into a training set without a trace, authors lose any leverage to negotiate or object. The industry may soon need receipts for corpus, the way software supply chains now demand a bill of materials, so that what was scanned and destroyed can at least be accounted for after the fact. Without such receipts, the default will remain that the largest buyers of books also become the least visible collectors of knowledge. The episode is a reminder that the cheapest data is often the one nobody knows was taken. Transparency, not scale, is what the public will ultimately demand from these pipelines.