Cultural barbarism: How AI companies are destroying the world’s books
source: https://www.telegraph.co.uk/news/2026/07/30/why-ai-companies-are-destroying-millions-of-old-books/
In its attempt to boost AI’s power, Silicon Valley is buying millions of rare editions, scanning them and then shredding the originals
Time almost stands still in Pieter de Vries’s antiquarian map and book shop in the historic centre of Haarlem. The 60-year-old Dutchman specialises in maps, illustrations and prints, many dating back to the 16th and 17th centuries. He also boasts a flourishing sideline in rare books.
Among his collection is a book on the history of Rome, published in 1558; travel books from the 17th, 18th and 19th centuries; precious first editions; books on natural history; and, until recently, a limited edition of The Highgrove Florilegium – a stunning two-volume illustrated guide to the diverse plant life grown in the gardens of the King’s beloved Highgrove House in Gloucestershire, signed by the then Prince Charles.
On an average day, the dealer will make one book sale. Footfall in the store never exceeds a slow trickle. Once in a while, the telephone will ring. Which is what made the email De Vries received in mid-June, containing a list of around 3,000 titles, so unusual.
The sender identified herself as “Nataly” from a company called 2077AI. “I’m taking part in a new project focused on collecting books in multiple languages, currently mostly in English,” Nataly wrote. “We have put together a very large list of editions we are currently trying to source (the list is attached), and are planning on making a fairly large order.”
De Vries didn’t dare open the attached spreadsheet because he assumed it was a scam. “In the rare book trade, it’s very seldom that people want to buy more than one book,” he says. “So if somebody comes and says, ‘I want a couple of hundred of your books,’ it’s very strange.”
It wasn’t a scam, however; it was something even more disturbing. It turned out to be part of the latest extraordinary attempt by Silicon Valley to boost the power of its AI machines. De Vries soon discovered that he was just one of hundreds of rare book dealers around the world who, this year, have been approached by companies hoovering up the globe’s supply of second-hand books on behalf of California’s tech giants.
The books are then fed into high-speed scanning machines that cut the spines off and shred or pulp the originals. The secret project, which was exposed by tech news site 404 Media this week, has been set up to train new AI models. But social media users have responded with horror at the scale of the operation.
According to 404 Media, one service called ISBNdb offered the ability to order up to one million books at a time, including rare books with almost no surviving copies. As one X user put it this week: “Books that survived wars, fires and centuries of handling are being shredded so an AI can learn to write a better marketing email.”
Elon Musk also voiced his disquiet. “I’ve asked the SpaceX AI team to preserve any rare books in a library and scan them the hard way rather than just cutting off the spine and scanning,” he said in an X post.
The seeds of the project were sown two years ago. In February 2024, Anthropic, the maker of Claude, one of the world’s most sophisticated AI chatbots, hired Tom Turvey, a former executive at Google Books, with a simple but staggering goal: obtain “all the books in the world”. Anthropic wanted to build up the knowledge of its AI model, and all the books in the world could become useful fodder to sate its appetite for data.
Marçal Font i Espí, a second-hand bookseller based in Badalona, a few miles outside Barcelona, and a literature teacher at the University of Barcelona, believes AI firms have reached “what’s called a data wall”: as the web becomes overrun with AI-generated “slop”, model developers are running short of fresh, human-made data. They still need “human-produced stuff”, he says, to improve.
Before moving to Anthropic, Turvey had worked as head of partnerships at Google’s own book-scanning project – but his new company’s books programme was a little different. He was asked to acquire the world’s supply of old books with as little “legal/practice/business slog” (as Anthropic’s internal communications described it) as possible.
A California judge’s decision in a court case looking at whether Anthropic owed money to authors whose books it used to train its models was released last month. The documents showed that, shortly after taking up his role, Turvey sent a few cursory emails to publishers about licensing books for AI training purposes. He quickly dropped that approach, though, and instead went directly to book distributors and retailers, looking to bulk-buy hard copies.
In all, Anthropic spent millions of dollars buying books en masse, then destroying them by ripping pages from their covers using a hydraulic-powered cutting machine, before throwing them away. The books didn’t have to be destroyed for the scanners to read them, but their destruction strengthened Anthropic’s legal case. The judge ruled that stripping books of their contents through digital scanning, then destroying them entirely, was not a breach of copyright because the digital version “replaced” the physical one: it was seen as an act of transformation rather than duplication, and thus not against copyright law.
Dario Amodei, chief executive of Anthropic (far left), is sworn in to testify at a US Senate hearing on the oversight of AI, July 2023 Credit: Saul Loeb
The oddity of that American law has its roots in older battles over book digitalisation. In essence, previous rights holders so stridently defended their copyright that AI companies must destroy books outright to use them.
The California ruling has created a financial incentive for AI firms to buy second-hand or unwanted book stock rather than negotiate licences with publishers, which – beyond the destruction of the books themselves – has copyright campaigners up in arms. “Yes, it may well be legal at the moment, but a lot of people think it shouldn’t be,” says Ed Newton-Rex, founder of Fairly Trained, a creator-rights group and outspoken critic of the AI industry’s approach to copyrighted work. He used to work in AI himself and argues that purchasing a physical copy should not automatically grant permission to use its contents to build a commercial model.
“If you are just going and spending $1 on a used book, with all of the money going to a book wholesaler, should that give you the right to train a commercial generative AI model on that book, which will then be able to compete with the author who wrote it?” Newton-Rex suggests that “a lot of people, myself included, think it shouldn’t.” Anthropic declined to comment.
Meanwhile, the act of destroying the physical books is also unpalatable to many, a fact of which AI companies seem to be aware: the cottage industry which has sprung up to support the mass purchase of old books is one shrouded in secrecy. As ISBNdb wrote in a blog post on its website: “The optics problem is real. ‘AI company destroys two million books’ is not a headline that generates sympathy.” (An ISBNdb spokesman admitted to The Telegraph it offered the service, but said it was a test “to explore demand for a service we never brought to life”.)
Font i Espí first became suspicious when his bookshop received bursts of unusual purchases, many involving postage costs far greater than the value of the books. There is no suggestion the orders came from ISBNdb. Font i Espí says booksellers in Germany, France, Spain and New Zealand have seen similar activity.
Newton-Rex says the secrecy is revealing. “Clearly both the provider of these books and the AI companies know that this is a terrible look and they don’t want the specifics to get out,” he says.
Court cases mentioning the scanning and destruction of books don’t mention specific titles other than those brought by the plaintiffs in the case, but 404 Media alleges that many have been rare. De Vries says none of the Dutch antiquarian dealers he knows have responded, but he believes less-specialised second-hand sellers may have done so. “There are a lot of booksellers who do second-hand books – just tonnes of them – and I think that they shipped quite a lot,” he says.
For Font i Espí, the danger is particularly acute when companies acquire the last accessible copy of an obscure work.
Copyright law makes little distinction between a battered Harry Potter paperback, one of millions printed, and one of only five surviving copies of an obscure antiquarian work. “Copyright law is very unconcerned with originals and physical artefacts,” says James Grimmelmann, a professor of digital and information law at Cornell University. Whether a book was printed in the millions or in an edition of five, “it treats them all as a copy”.
Font i Espí stresses that he is not against digitisation itself. “I’m not against AI. I understand that books and old knowledge must be transferred to new formats,” he says. But the goal should be to “digitalise it properly, not destroy the original document.”
For Newton-Rex, though, whether the books are rare or plentiful in print doesn’t matter. “There is surely no more fitting image in the generative AI age for the exploitation that underlies this technology than almost trillion-dollar companies buying books for a few cents or a dollar each, scanning them, training on them, then destroying them, essentially subsuming culture.”
And, legal as it may currently be, not everyone in Silicon Valley is content with treating books in this way. “There’s something particularly misanthropic about the mechanised destruction of such intimate human objects,” wrote Matan Grinberg, chief executive of Factory AI, an AI company. “History seldom looks kindly on those who destroy books, whatever the reasons.”
Or, as bookseller De Vries puts it: “It is a kind of barbarism… like book burning.”