Skip to main content

The silent plunder · Investigation by Xavier Vinaixa & Marçal Font Espí

AI companies are buying and destroying books

How a mysterious company buys second-hand books from bookstores across half the world to train AI models — then physically destroys the copies.

This page collects the joint investigation conducted by Xavier Vinaixa (xaviviro) and UB professor, bookseller and poet Marçal Font Espí from the Llibreria Fènix bookshop in Badalona, after spotting anomalous orders from a foreign company systematically buying Catalan non-fiction — often copies that had been sitting in the back room for years.

What looked like an unusual buyer turned out to be an industrial chain: the firms buy the books, ship them to a plant where the spines are cut, the pages are scanned and shredded, and what is left of the volumes is pulped. All of it to harvest training data — human-written text to train artificial intelligence models before the web is overrun by AI-generated content, what specialists call the «data wall».

What we uncovered

You may have seen the July 2026 wave of coverage about ISBNdb and AI companies shredding rare books for training data. That's the same phenomenon — but that reporting stops at the US middleman. We traced the European end months earlier: the ZoomBooks orders, the PrepFort logistics chain in Illinois, and the destructive scanning behind Anthropic's Project Panama. And unlike that coverage, we propose a way out.

An international network silently acquires lots of rare and out-of-print books of low commercial value from second-hand bookshops. They are chosen precisely because they are human texts barely present on the internet — no digital footprint, no editorial scrutiny, no effective copyright protection in practice. Anything printed before 2022 is clean human prose, free of AI slop, the scarce raw material labs need to keep their models from degrading into what researchers call model collapse.

The destruction is not incidental, it is what makes the practice defensible in court: under the US doctrine of fair use, Judge Alsup's June 2025 ruling in the Anthropic case treated training on lawfully bought books as legitimate provided the physical copy is destroyed once it has been scanned. So the shredder is a legal requirement, not a side effect. The consequence is not only economic for authors and publishers: it is cultural. Minoritised languages such as Catalan lose written heritage not through human neglect but through algorithmic voracity. Our investigation documents the mechanism and proposes data sovereignty pathways — protection and licensing instead of destruction — in the Substack trilogy.

Frequently asked questions

Are AI companies really buying and destroying books?
Yes. Companies acting on behalf of AI labs buy rare and out-of-print second-hand books from bookshops, ship them to industrial facilities where the spines are cut off and the pages scanned, and the physical copies are destroyed. We documented the European end of the chain from the Llibreria Fènix in Badalona: the ZoomBooks orders and the PrepFort logistics chain in Illinois. The case was later picked up by The Telegraph, El País, BBC News Mundo and dozens of other outlets.
Which AI companies buy second-hand books?
The labs rarely buy directly — the orders arrive through intermediaries, which is what makes the chain hard to trace. Our investigation followed purchases placed by ZoomBooks, a Canadian company, routed through the PrepFort logistics chain in Illinois; ZoomBooks publicly denied the accusations published by Demócrata. Destructive scanning also underpins Anthropic's «Project Panama», revealed by the Washington Post.
Why do AI companies want old, rare and out-of-print books?
Because they are the last clean training data left. Anything printed before 2022 is human prose free of AI slop, and rare or out-of-print titles have almost no digital footprint, so they are not already inside the training sets. That scarcity is what keeps models from degrading into what researchers call model collapse.
What is the «data wall»?
The point at which AI labs run out of freely available digitised text on the internet and have to find new sources to keep scaling their models. It is the economic reason behind the sudden interest in second-hand bookshops: once the free knowledge of the internet was exhausted, the labs went looking for what had never been digitised in the first place.
Why are the books destroyed instead of resold?
Because the destruction is what makes the practice defensible in court. Under the US doctrine of fair use, Judge Alsup's June 2025 ruling in the Anthropic case treated training on lawfully bought books as legitimate provided the physical copy is destroyed once it has been scanned. The shredder is a legal requirement, not a side effect.
Can a bookseller stop their books being used to train AI?
No legal mechanism prevents a lawful sale from being followed by destructive scanning. What booksellers can do in practice is spot the pattern: several small orders from the same buyer minutes apart, for out-of-print non-fiction with no commercial demand. The structural answer we propose is Cedulari — cataloguing and licensing written heritage instead of letting it be destroyed.
What is Anthropic's «Project Panama»?
The internal name of the book-acquisition and scanning programme revealed by the Washington Post, under which Anthropic obtained physical books at scale, cut the spines off, scanned the pages and destroyed what was left. It is the clearest documented case of an AI lab industrialising destructive scanning, and it runs on the same logic our investigation traced from the other end of the chain: buy lawfully, scan, destroy.
What is ZoomBooks?
The Canadian company whose orders first exposed the pattern to us. From the Llibreria Fènix in Badalona we tracked bursts of small purchases of out-of-print Catalan non-fiction, routed through the PrepFort logistics chain in Illinois, with the final recipient never named. ZoomBooks publicly denied the accusations published by Demócrata. The tell was economic rather than bibliographic: postage repeatedly cost more than the books themselves, which makes no sense for a reseller and perfect sense for a buyer who only wants the text.
What is destructive scanning?
The industrial method for turning a printed book into training data as fast as possible: the spine is cut off with a guillotine or hydraulic cutter, the loose pages are fed through a production sheet-feed scanner, and the paper is then pulped. It is far quicker and cheaper than scanning a bound volume page by page, and — since Judge Alsup's ruling — the destruction is not a regrettable by-product but the step that makes the copy legally defensible.
How widespread is the practice?
It is not a local anomaly. Beyond the Catalan bookshops where we first documented it, The Telegraph confirmed the same ordering pattern with booksellers in Germany, France, Spain and New Zealand, and reported a Haarlem antiquarian dealer who received a list of some 3,000 titles from a company called 2077AI and assumed it was a scam. 404 Media separately revealed the use of the ISBNdb catalogue to target acquisitions. The buyers change; the pattern — bulk orders of obscure, un-digitised titles — does not.
Is Anthropic buying and destroying books?
Yes, and it has been established in court rather than merely alleged. Anthropic hired Tom Turvey in February 2024 with the brief of obtaining «all the books in the world», and Judge Alsup's June 2025 ruling in Bartz v. Anthropic accepted training on lawfully purchased, destroyed books as fair use while treating pirated copies very differently — the piracy claims ended in a 1.5 billion dollar settlement, finally approved in July 2026 at roughly 3,100 dollars per work across 440,490 claims. Buying and shredding a book is free; downloading it costs billions. The incentive is exactly inverted.
What is «AI slop» and why does it make old books valuable?
«Slop» is the mass of low-quality, AI-generated text now flooding the open web. It contaminates any dataset scraped after roughly 2022, and models trained on their own output degrade — what researchers call model collapse. Paper printed before that date is verifiably human, and an out-of-print book that was never digitised is text no competitor already holds. That combination — human-written and absent from the internet — is what makes a worthless second-hand book a strategic asset.
Which books are AI companies buying?
Not the valuable ones. The orders concentrate on out-of-print non-fiction with no commercial demand: local histories, technical manuals, regional cookbooks, specialist monographs, minority-language titles — the stock that sits in a back room for years. Their bibliographic obscurity is the point, and it is also what makes the loss irreversible: when the copy destroyed is one of the last accessible ones, the book survives only inside a private training set, where it can no longer be consulted, leafed through or checked.
Who first uncovered that AI companies were destroying books?
The European end of the chain was traced by Marçal Font Espí, bookseller at the Llibreria Fènix and University of Barcelona literature teacher, together with Xavier Vinaixa Roselló, CTO at Sorensen.ai — months before the story broke in the United States. Their series «The silent plunder» documented the buying pattern, the logistics route and the legal incentive, and is the origin the later coverage rests on, from The Telegraph and BBC News Mundo to Libération and El País.

Cedulari · Our proposal

The solution we propose: Cedulari

We don't stop at denouncing the problem. For over a year we have been building a constructive answer: Cedulari, the documented memory of Catalan publishing. It is both a database of authors, books and editions —focused on nearly a century of Catalan publishing, from the 1920s to the adoption of the ISBN, a heritage that today is recorded nowhere— and a set of tools that give AI models provenance, traceability and verifiable sources instead of the usual hallucinations.

It was precisely this knowledge —of the second-hand book trade, the print runs, the pseudonyms and the clandestine editions held by only five or six antiquarian booksellers in the country— that let us uncover the silent plunder. Cedulari exists to preserve that heritage and make it available to AI lawfully and traceably. The project, led by Marçal Font, Xavi Vinaixa and Sorensen.ai with advice from the University of Barcelona's literary studies department, is now seeking funding.

Discover Cedulari.cat

Who investigated

01

Xavier Vinaixa Roselló

AI expert, CTO at Sorensen.ai. Pulled the technical thread: identifying the purchase pattern, analysing the logic of the «data wall» and framing it within the wider debate on the real cost of training large language models.

02

Marçal Font Espí

Professor in the Department of Literary Studies at the UB, poet and bookseller, owner of the Llibreria Fènix bookshop in Badalona. Spotted the anomalous orders, documented the requests and contributed first-hand knowledge of the second-hand book trade.

The series: «The silent plunder»

The four instalments of the series are co-authored by Xavier Vinaixa and Marçal Font Espí, published on xaviviro.substack.com, with a companion piece on Marçal Font's Substack (maralfontesp.substack.com):

  1. The silent plunder

    xaviviro.substack.com →

  2. The silent plunder II: the erased confession

    xaviviro.substack.com →

  3. The silent plunder III: protect and license

    xaviviro.substack.com →

  4. The silent plunder IV: the architecture of evasion

    xaviviro.substack.com · 2026-06-25 →

  5. Companion piece

    And we'll make drums out of the books

    maralfontesp.substack.com · 2026-05-18 →

RTVE · TV3 · 3Cat · 2026

On television

The viral 3Cat Info reel

700k+ views

3Cat Info's (TV3) report on the literary plunder went viral on Instagram, with over 700,000 views. It features Marçal Font (Llibreria Fènix) and Xavier Vinaixa explaining how some companies buy and destroy books to train artificial intelligence.

@3catinfo · TV3 · Instagram →

«Cafè d'idees» (RTVE) — 15 June 2026

RTVE's «Cafè d'idees», hosted by Gemma Nierga, welcomes Xavier Vinaixa (technical director of Sorensen) and bookseller Marçal Font (Llibreria Fènix in Badalona) to explain how some companies buy books, train artificial intelligence and then destroy them —«they've hit a ceiling, there's no more data»— and to call for European data sovereignty in the face of the mass purchase of books to feed AI.

«Cafè d'idees» · RTVE · 15 June 2026 · rtve.es →

«Tot es mou» — 27 May 2026

TV3's flagship news magazine «Tot es mou» invites Marçal Font (Llibreria Fènix bookshop in Badalona) and Xavier Vinaixa (CTO at Sorensen) on the studio set to explain how they spotted the anomalous purchases, how they pulled the thread to Anthropic and the story originally broken by the Washington Post, and why preserving the physical document is a matter of cultural sovereignty — not just economics. Programme aired in Catalan.

«Tot es mou» · TV3 · 27 May 2026 · 3cat.cat →

Telenotícies (TV3 newscast) — 31 May 2026

TV3's flagship newscast Telenotícies ran a piece on the case: booksellers denounce the destruction of books to train AI models and call on the Ministry of Culture to step in to protect literary heritage. The report picks up the case uncovered by the investigation. Aired in Catalan.

Telenotícies · TV3 · 31 May 2026 · 3cat.cat →

Press coverage

The investigation has been picked up by television, the press, national radio and international outlets:

Read the investigation

Read on Substack