ActaVerum.
// AI · ANALYSIS

Why AI wants your used books

Used bookshops in the Netherlands, Australia and the UK report bulk orders from buyers who will not explain where the books end up. What is proven is the Anthropic case: millions of copies bought, scanned and destroyed.

By Newsroom·Aug 20, 2026·AI
crowded bookshop shelves filled with used books
Illustrative photo: the shelves of an old bookshop in Bilbao, used here to represent the secondhand stock at the center of the bulk-order reports. Iñaki del Olmo / Unsplash

Pieter de Vries runs an antiquarian bookshop in Haarlem, in the Netherlands. One day an email arrived from someone called Natalia, at a company called 2077AI, describing a project to collect books in several languages. Attached was a spreadsheet of 3,001 ISBNs: academic titles from 2020 and 2021 published by Emerald, Elsevier, Wiley, Routledge and Oxford University Press, covering business, education, engineering, public policy and medicine. Shipping address in China. He assumed it was phishing and ignored it.¹

Other Dutch dealers got the same request and made the same call. Similar reports followed from Germany and Switzerland.¹ In August, the Guardian ran two rounds of reporting, one on Australian secondhand dealers and one on shops in the UK and Ireland.² ³ A British seller said he had shipped roughly 6,000 books since January. David Gower-Spence, of BookLovers of Bath, sent about 200 between May and July. One recent order bundled an Estonian translation of John le Carré's The Mission Song, a particular Anne Brontë imprint, and the October 1983 issue of Warship

None of these buyers asked for a bulk discount. Orders under several different names went to the same addresses, and one postcode traced back to freight warehouses near Heathrow.³ On the other side of the world, in Benalla, in rural Victoria, Delfina Manor shipped three boxes in May to the Canadian firm Zoom Books; John Sainsbury, in Melbourne, filled orders that mixed engineering, cycling, poetry and suburban local history with no collector's logic behind them.²

The suspicion is the same everywhere, and the booksellers state it plainly: the books are headed for destructive scanning, to become AI training data.

What the record already shows

It's worth separating what sits in a document from what is still a hunch, because the gap here is wide.

On June 23, 2025, Judge William Alsup, of the US District Court for the Northern District of California, issued his fair use ruling in Bartz v. Anthropic.⁴ The opinion describes what Anthropic did, drawn from the case record. The company downloaded more than seven million pirated copies of books: Books3, with 196,640 titles, in 2021; at least five million from Library Genesis in June 2021; at least two million from the Pirate Library Mirror in July 2022.

Then the strategy changed. In February 2024, Anthropic hired Tom Turvey, formerly head of partnerships for Google's book-scanning project, and tasked him with obtaining "all the books in the world". Turvey let the publisher conversations wither and took a different route. In the ruling's own words, the company "spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form, discarding the paper originals."⁴

The Washington Post, which reviewed more than 4,000 pages of case documents unsealed in January 2026, reported the operation's internal name: Project Panama. A planning document set out the goal as an effort "to destructively scan all the books in the world", and an April 2024 memo asked for discretion, since "we don't want it to be known that we are working on this".⁵ ⁶ Vendor proposals discussed 500,000 to two million books over six months.⁵

Alsup held that converting print to digital was fair use: for each purchased print copy it destroyed, the company kept an internal digital replacement that was easier to store and search.⁴ The library built from pirated copies was another matter. That half ended in a $1.5 billion settlement, about $3,000 per work across roughly 482,000 books, granted final approval on July 20, 2026 by Judge Araceli Martínez-Olguín.⁷

What is still guesswork

Nobody has shown that the 2026 orders come from AI companies, and nobody has produced a rare book that went into a shredder.

Alex Reisner, at The Atlantic, traced the orders and landed on Zoom Books as the largest recurring buyer, with some deliveries routed through PrepFort, a third-party logistics operator. The profile of the books also cuts against the version that spread online: mostly nonfiction from the 1970s to the 1990s, the kind of title that ages badly and falls out of print, rather than anything a collector would chase.⁸ Charlie Becker, of Becker's Books in Houston, said 95 of his last 100 sales on one platform went to a single buyer.⁸

Zoom Books describes itself as a book recycler, says it resells copies intact and does not digitize them, and cites commercial confidentiality when asked whether it resells to AI firms.² Anthropic says it has never bought from Zoom Books, that it buys through conventional commercial channels, and that "none of our data acquisition programs buy and destroy rare or antiquarian books".² In August it added that "sourcing books is a widely used approach for training large language models across the AI industry".³

Snopes, in a fact check published on August 3, rated the viral version mostly true and flagged where it overreaches: an email from Turvey did say that "less common books are a great place to start", but that points at hard-to-find professional and academic references, not antiquarian first editions.⁶ Anthropic is the only company on the record. For the others, the public evidence runs in neither direction.

Why paper is suddenly worth this much

The clearest commercial argument came from the party trying to sell the service.

On July 21, 404 Media reported that ISBNdb, which runs one of the largest bibliographic databases anywhere, had a live page offering bulk print-book acquisition to AI labs, from a thousand to a million copies per engagement, under NDAs that kept the client's identity out of sight. The copy sold books as "curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative." It also named the discomfort: "the optics problem is real", because "'AI company destroys two million books' is not a headline that generates sympathy".⁹

Nine days later the page came down, and ISBNdb said it had "never purchased, scanned, or sold a book, for AI training or anything else", calling the pages "a test of market interest; no such service was ever brought to life".¹⁰ As of August 17, both of those addresses redirect to the company's homepage.

Underneath the sales pitch sit two real technical worries. The first is model collapse: train a model over and over on text that other models produced and the distribution drifts, low-frequency events vanish and output variety shrinks, and fresh human data has to be fed back in periodically to keep the effect in check.¹¹ Print published before 2022 predates the flood of synthetic text, and that guarantee is hard to buy on the open web.

The second is data poisoning. In an experiment by Anthropic with the UK AI Security Institute and the Alan Turing Institute, roughly 250 manipulated documents were enough to backdoor models from 600 million to 13 billion parameters, without scaling in proportion to model size.¹² The study says nothing about books and does not show that paper is immune; that bridge belongs to the vendor, not to the authors.

There is also the obvious point: a large share of what exists in print never reached the internet at all. Out-of-print technical manuals, dissertations, regional monographs, translations into small languages. It is precisely the material a model cannot find on its own.

Cutting the spine is a cost decision

Non-destructive scanning exists, and it is what libraries use, because the copy has to go back on the shelf. Slicing the spine is faster and cheaper: once the pages are loose, a sheet feeder does the work a human would otherwise do page by page.

The consequence is that the book leaves circulation and the surviving copy is locked away. Alsup's ruling is explicit on this: Anthropic's digital library was internal, "not for sharing nor sale outside the company".⁴ A copy that sat in a shop and could be sold, lent, quoted and resold becomes a PDF no outsider will ever open. Calling that the destruction of knowledge overstates it, since the text usually survives in other copies. What happens is narrower: an object leaves the public world and carries on as a private asset.

The economics are lopsided. The licensed route has a published price: in November 2024, HarperCollins struck a deal with an undisclosed technology company, reported to be Microsoft, to license backlist nonfiction at $5,000 per title for three years, half of it to the author, on an opt-in basis.¹³ ¹⁴ A used copy on a marketplace costs a few dollars and requires no negotiation, no consent and no term limit. Alsup's ruling left both routes open, and one of them is far cheaper.

Brazil hasn't written the rule yet

Brazil has no specific rule for AI training yet. The version of PL 2338/2023 approved by the Senate in December 2024 carries both halves of this argument: a limited exception for text and data mining and, in article 65, a remuneration right for holders of works used in training. As of August 17, 2026, the bill was awaiting the rapporteur's opinion in a special committee of the Chamber of Deputies.¹⁵

In the meantime, a Brazilian secondhand bookshop selling into an international marketplace is in the same position as the Dutch one. An order comes in, it gets packed, it ships, and where the book ends up is anyone's guess.

What the booksellers are saying

What stands out is how little the trade's objection has to do with copyright.

Authors and publishers whose works are covered by the $1.5 billion settlement are entitled to compensation for the piracy. The booksellers are not parties to that case, report no obvious financial loss and, in practice, have had a strong sales season: the buyers pay full price and clear dead stock. Their complaint runs elsewhere. Stuart Manley, of Barter Books in Alnwick, says the strangest part is that the orders follow no theme at all, a pattern he describes as sporadic and presumably machine-driven, unlike any customer he recognizes.³ Jim Shaughnessy, of MW Books in Ireland, and Gower-Spence describe the same sense of supplying something that will not identify itself.³

What grates is the asymmetry of information. The seller does not know what is being bought, by whom, or for what. And none of it is recorded anywhere: no agency tracks this market, and no list says which titles went through the blades. If an irreplaceable copy does go in, nobody will find out, because there is nowhere to look.

Verdict

The story has three layers, and most of the confusion comes from treating them as one.

An AI company bought millions of used books, cut off the bindings and threw the paper away. That is in the case record, corroborated by reporting on internal documents, and disputed by no one.⁴ ⁵ A US federal court held in this case that replacing each purchased copy with an internal digital copy is fair use. The ruling carries weight in the debate, but does not by itself bind every other US court.⁴

The strange 2026 orders are a real phenomenon, documented by several newspapers across three continents, with named buyers, one of whom declines to discuss where the books go.² ³ ⁸ Connecting those boxes to a shredder, though, is still inference rather than observation.

And "rare books are being destroyed" is the weakest link in the chain. What the reporting actually describes is ordinary, out-of-print, unglamorous stock, the useful clutter of a secondhand shop. The structure of the market deserves more attention than the lost-treasure scenario: an anonymous, lawful supply chain with no record of what goes into it. A rare book usually has a catalogue, a provenance and someone who would notice it missing. An out-of-print soil-mechanics manual may be catalogued too, but rarely has anyone tracking the fate of each copy. That is exactly the material leaving circulation by the boxful.²

Sources

  1. This Dutch bookseller thought a request for 3,000 copies was 'spam or phishing.' · Fortune (Tatiana Sataua) · https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/ · 2026-07-31.
  2. 'More than just objects': Australian book sellers raise alarm over 'horrific' destruction of rare titles to feed AI · The Guardian · https://www.theguardian.com/technology/2026/aug/02/australian-book-sellers-alarm-destruction-rare-titles-ai-supply-chain · 2026-08-02.
Show 13 more sourcesHide sources
  1. Secondhand booksellers in UK and Ireland suspect AI firms behind 'strange' bulk orders · The Guardian (Dan Milmo), via AOL · https://www.aol.co.uk/articles/secondhand-booksellers-uk-ireland-suspect-080051000.html · 2026-08-15.
  2. Bartz v. Anthropic PBC, Order on Fair Use · US District Court, Northern District of California, No. C 24-05417 WHA, doc. 231 (Judge William Alsup) · https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/ANTHROPIC%20fair%20use.pdf · 2025-06-23.
  3. Anthropic 'destructively' scanned millions of books to build Claude · The Washington Post · https://www.washingtonpost.com/technology/2026/01/27/anthropic-ai-scan-destroy-books/ · 2026-01-27.
  4. Are AI companies scanning and destroying millions of books, including rare titles? · Snopes (Anna Rascouët-Paz) · https://www.snopes.com/fact-check/ai-companies-destroying-rare-books/ · 2026-08-03.
  5. Anthropic's landmark $1.5B copyright settlement is approved · TechCrunch (Kirsten Korosec) · https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved/ · 2026-07-20.
  6. Someone Is Mysteriously Snapping Up Used Books Around the World · The Atlantic (Alex Reisner) · https://www.theatlantic.com/technology/2026/08/ai-companies-buying-used-books-for-data/688167/ · 2026-08.
  7. AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop · 404 Media (Emanuel Maiberg) · https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/ · 2026-07-21.
  8. Company Offering Printed Books to Train AI Stops After 404 Media Coverage · 404 Media (Samantha Cole) · https://www.404media.co/ai-company-training-scanning-books-database-isbndb/ · 2026-07-31.
  9. Shumailov, I. et al. AI models collapse when trained on recursively generated data · Nature 631, 755–759 · DOI 10.1038/s41586-024-07566-y · 2024.
  10. A small number of samples can poison LLMs of any size · Anthropic, UK AI Security Institute and the Alan Turing Institute · https://www.anthropic.com/research/small-samples-poison · 2025-10-09.
  11. HarperCollins AI Licensing Deal · The Authors Guild · https://authorsguild.org/news/harpercollins-ai-licensing-deal/ · 2024.
  12. Agents, Authors Question HarperCollins AI Deal · Publishers Weekly · https://www.publishersweekly.com/pw/by-topic/industry-news/publisher-news/article/96533-agents-authors-question-harpercollins-ai-deal.html · 2024.
  13. PL 2338/2023, legislative tracking · Brazilian Chamber of Deputies · https://www.camara.leg.br/proposicoesWeb/fichadetramitacao?idProposicao=2487262 · consulted 2026-08-17.

Comments 0