Skip to main content

The silent plunder · Investigation by Xavier Vinaixa & Marçal Font Espí

AI companies are buying and destroying books

How a mysterious company buys second-hand books from bookstores across half the world to train AI models — then physically destroys the copies.

This page collects the joint investigation conducted by Xavier Vinaixa (xaviviro) and UB professor, bookseller and poet Marçal Font Espí from the Llibreria Fènix bookshop in Badalona, after spotting anomalous orders from a foreign company systematically buying Catalan non-fiction — often copies that had been sitting in the back room for years.

What looked like an unusual buyer turned out to be an industrial chain: the firms buy the books, ship them to a plant where the spines are cut, the pages are scanned and shredded, and what is left of the volumes is pulped. All of it to harvest training data — human-written text to train artificial intelligence models before the web is overrun by AI-generated content, what specialists call the «data wall».

What we uncovered

You may have seen the July 2026 wave of coverage about ISBNdb and AI companies shredding rare books for training data. That's the same phenomenon — but that reporting stops at the US middleman. We traced the European end months earlier: the ZoomBooks orders, the PrepFort logistics chain in Illinois, and the destructive scanning behind Anthropic's Project Panama. And unlike that coverage, we propose a way out.

An international network silently acquires lots of rare and out-of-print books of low commercial value from second-hand bookshops. They are chosen precisely because they are human texts barely present on the internet — no digital footprint, no editorial scrutiny, no effective copyright protection in practice. Anything printed before 2022 is clean human prose, free of AI slop, the scarce raw material labs need to keep their models from degrading into what researchers call model collapse.

The destruction is not incidental, it is what makes the practice defensible in court: under the US doctrine of fair use, Judge Alsup's June 2025 ruling in the Anthropic case treated training on lawfully bought books as legitimate provided the physical copy is destroyed once it has been scanned. So the shredder is a legal requirement, not a side effect. The consequence is not only economic for authors and publishers: it is cultural. Minoritised languages such as Catalan lose written heritage not through human neglect but through algorithmic voracity. Our investigation documents the mechanism and proposes data sovereignty pathways — protection and licensing instead of destruction — in the Substack trilogy.

Cedulari · Our proposal

The solution we propose: Cedulari

We don't stop at denouncing the problem. For over a year we have been building a constructive answer: Cedulari, the documented memory of Catalan publishing. It is both a database of authors, books and editions —focused on nearly a century of Catalan publishing, from the 1920s to the adoption of the ISBN, a heritage that today is recorded nowhere— and a set of tools that give AI models provenance, traceability and verifiable sources instead of the usual hallucinations.

It was precisely this knowledge —of the second-hand book trade, the print runs, the pseudonyms and the clandestine editions held by only five or six antiquarian booksellers in the country— that let us uncover the silent plunder. Cedulari exists to preserve that heritage and make it available to AI lawfully and traceably. The project, led by Marçal Font, Xavi Vinaixa and Sorensen.ai with advice from the University of Barcelona's literary studies department, is now seeking funding.

Discover Cedulari.cat

Who investigated

01

Xavier Vinaixa Roselló

AI expert, CTO at Sorensen.ai. Pulled the technical thread: identifying the purchase pattern, analysing the logic of the «data wall» and framing it within the wider debate on the real cost of training large language models.

02

Marçal Font Espí

Professor in the Department of Literary Studies at the UB, poet and bookseller, owner of the Llibreria Fènix bookshop in Badalona. Spotted the anomalous orders, documented the requests and contributed first-hand knowledge of the second-hand book trade.

The series: «The silent plunder»

The four instalments of the series are co-authored by Xavier Vinaixa and Marçal Font Espí, published on xaviviro.substack.com, with a companion piece on Marçal Font's Substack (maralfontesp.substack.com):

  1. The silent plunder

    xaviviro.substack.com →

  2. The silent plunder II: the erased confession

    xaviviro.substack.com →

  3. The silent plunder III: protect and license

    xaviviro.substack.com →

  4. The silent plunder IV: the architecture of evasion

    xaviviro.substack.com · 2026-06-25 →

  5. Companion piece

    And we'll make drums out of the books

    maralfontesp.substack.com · 2026-05-18 →

RTVE · TV3 · 3Cat · 2026

On television

The viral 3Cat Info reel

700k+ views

3Cat Info's (TV3) report on the literary plunder went viral on Instagram, with over 700,000 views. It features Marçal Font (Llibreria Fènix) and Xavier Vinaixa explaining how some companies buy and destroy books to train artificial intelligence.

@3catinfo · TV3 · Instagram →

«Cafè d'idees» (RTVE) — 15 June 2026

RTVE's «Cafè d'idees», hosted by Gemma Nierga, welcomes Xavier Vinaixa (technical director of Sorensen) and bookseller Marçal Font (Llibreria Fènix in Badalona) to explain how some companies buy books, train artificial intelligence and then destroy them —«they've hit a ceiling, there's no more data»— and to call for European data sovereignty in the face of the mass purchase of books to feed AI.

«Cafè d'idees» · RTVE · 15 June 2026 · rtve.es →

«Tot es mou» — 27 May 2026

TV3's flagship news magazine «Tot es mou» invites Marçal Font (Llibreria Fènix bookshop in Badalona) and Xavier Vinaixa (CTO at Sorensen) on the studio set to explain how they spotted the anomalous purchases, how they pulled the thread to Anthropic and the story originally broken by the Washington Post, and why preserving the physical document is a matter of cultural sovereignty — not just economics. Programme aired in Catalan.

«Tot es mou» · TV3 · 27 May 2026 · 3cat.cat →

Telenotícies (TV3 newscast) — 31 May 2026

TV3's flagship newscast Telenotícies ran a piece on the case: booksellers denounce the destruction of books to train AI models and call on the Ministry of Culture to step in to protect literary heritage. The report picks up the case uncovered by the investigation. Aired in Catalan.

Telenotícies · TV3 · 31 May 2026 · 3cat.cat →

Press coverage

The investigation has been picked up by television, the press, national radio and international outlets:

Frequently asked questions

Are AI companies really buying and destroying books?
Yes. Companies acting on behalf of AI labs buy rare and out-of-print second-hand books from bookshops, ship them to industrial facilities where the spines are cut off and the pages scanned, and the physical copies are destroyed. We documented the European end of the chain from the Llibreria Fènix in Badalona: the ZoomBooks orders and the PrepFort logistics chain in Illinois. The case was later picked up by The Telegraph, El País, BBC News Mundo and dozens of other outlets.
Which AI companies buy second-hand books?
The labs rarely buy directly — the orders arrive through intermediaries, which is what makes the chain hard to trace. Our investigation followed purchases placed by ZoomBooks, a Canadian company, routed through the PrepFort logistics chain in Illinois; ZoomBooks publicly denied the accusations published by Demócrata. Destructive scanning also underpins Anthropic's «Project Panama», revealed by the Washington Post.
Why do AI companies want old, rare and out-of-print books?
Because they are the last clean training data left. Anything printed before 2022 is human prose free of AI slop, and rare or out-of-print titles have almost no digital footprint, so they are not already inside the training sets. That scarcity is what keeps models from degrading into what researchers call model collapse.
What is the «data wall»?
The point at which AI labs run out of freely available digitised text on the internet and have to find new sources to keep scaling their models. It is the economic reason behind the sudden interest in second-hand bookshops: once the free knowledge of the internet was exhausted, the labs went looking for what had never been digitised in the first place.
Why are the books destroyed instead of resold?
Because the destruction is what makes the practice defensible in court. Under the US doctrine of fair use, Judge Alsup's June 2025 ruling in the Anthropic case treated training on lawfully bought books as legitimate provided the physical copy is destroyed once it has been scanned. The shredder is a legal requirement, not a side effect.
Can a bookseller stop their books being used to train AI?
No legal mechanism prevents a lawful sale from being followed by destructive scanning. What booksellers can do in practice is spot the pattern: several small orders from the same buyer minutes apart, for out-of-print non-fiction with no commercial demand. The structural answer we propose is Cedulari — cataloguing and licensing written heritage instead of letting it be destroyed.

Read the investigation

Read on Substack