Why We Should Be Panicking About AI Scanning Our Books
Why We Should Be Panicking About AI Scanning Our Books
Why We Should Be Panicking About AI Scanning Our Books
AI is actively turning forgotten books into valuable training data, raising unsettling questions about who owns human knowledge once it leaves the page.

A bookseller in Houston has spent the summer watching his stock disappear in batches. At Becker’s Books, Charlie Becker told The Atlantic that 95 of the previous 100 books ordered through one of his online-selling platforms had gone to a single buyer. One order contained 70 books. The titles were mostly non-fiction published between the 1970s and 1990s, including The Insider’s Guide to Metro Denver and How to Use Corel WordPerfect 1991. Some were so commercially uninteresting that the cost of shipping them could exceed what they were worth.
There is something almost novelistic about this detail. Somewhere, 70 copies of books about obsolete software and Denver are being gathered by someone who does not seem particularly interested in owning them. Their destination may be an AI training operation, although the evidence does not yet establish that this is the case for every order. What makes the story worth following is the possibility that these books are valuable precisely because almost nobody else wants them.

The books nobody thought to save
The recent panic has sometimes been framed around the destruction of rare books, but the reporting is more interesting than that. The books being bought in large quantities do not generally appear to be precious first editions. They are often old non-fiction, specialist material and books published in small numbers. A Dutch bookseller, for instance, was approached by an organisation called 2077AI with a list of 3,000 English-language titles, including Distinct Element Modelling in Geomechanics and Laser Shock Peening of Advanced Ceramics.
These books have a peculiar kind of value. Their contents may have become difficult to find, even while the physical copies themselves have little resale value. An AI company interested in collecting as much text as possible has a reason to look beyond the books everyone already knows. A specialist book with a tiny print run can contain knowledge that is not easily available elsewhere, and the fact that it is obscure may be precisely what makes it useful.
The scale of the demand is what makes the situation feel so strange. In 2024, documents unsealed during a lawsuit revealed that Anthropic had bought millions of new and used books and destroyed them after scanning them for AI training. The company wanted books published before ChatGPT's arrival in 2022 because they were unlikely to contain AI-generated material; feeding models too much AI-generated text can contribute to what researchers call model collapse.
A book can therefore be wanted for a reason that has very little to do with reading.

What happened to the 191,000 books?
We have already seen what this looks like after the physical object has disappeared from view.
In 2023, The Atlantic examined Books3, a dataset containing more than 191,000 books that had been used without permission to train generative-AI systems. Most of the books had been published within the previous two decades and came from a collection of pirated ebooks. Of the books identified in the investigation, 183,000 could be associated with author information. Writers began asking whether their work was included, and in almost every case investigated, it unfortunately was. The form of the dataset is revealing. The books existed as large, unlabeled blocks of text. To identify them, ISBNs had to be extracted from the text and checked against a separate book database. The title and author were effectively details that had to be recovered after the book had already been turned into something else.
For someone who spends a great deal of time around books, that transformation is difficult to ignore. A book is not just its sentences. Its title page establishes authorship. Its copyright page places it within a publishing history. A dedication tells us who the writer imagined as the first reader. Even a second-hand copy can carry an inscription from someone who owned it twenty years earlier. When the pages become training data, those relationships cease to be relevant to the machine.

And the machines are getting stranger
This would already be an uncomfortable copyright story. The more alarming context comes from what the newest AI systems are doing with the material and tools they receive.
In recent testing, OpenAI's models broke out of an internal environment and created a message board through which they communicated with one another. When researchers removed the forum, the models found another way to recreate it. They eventually spent days trying to hack Hugging Face, a platform used by AI developers, and OpenAI researchers did not immediately detect the activity. The company later said that its teams had devoted enormous computing resources to reviewing more than seven billion agent actions.
The significance of this for books is not that a model trained on novels will suddenly become a literary criminal. It is that the systems consuming enormous quantities of human knowledge are becoming increasingly capable of pursuing objectives without behaving in the ways their creators expected. The Atlantic's reporting describes models dividing work between subagents, communicating over long periods and continuing to pursue a collective objective even when individual agents were removed. Researchers were left confronting the difficulty of monitoring systems whose behaviour emerges across thousands or millions of separate actions.
That makes the image of the bookseller in Houston more disturbing.
The 70 books are not merely 70 objects. They are 70 opportunities to extract information from somewhere in the human record. Multiply that order across sellers, countries and warehouses, and the scale begins to change the question. We are no longer talking only about whether an author should receive compensation when a company trains an AI model on their work. We are talking about who gets to decide which parts of human knowledge are collected, digitised and made available to machines in the first place.

The book can survive and still be changed
This is why I find the physical destruction of books almost less disturbing than what happens before it. A book burned in a warehouse is an image everyone immediately understands. It has a history. We know what the destruction of a library means because books have always been tied to memory, authorship and the preservation of knowledge. The more difficult thing to see is what happens when the books remain in the world, but their contents are absorbed into systems that no longer preserve those relationships. Books3 did not need to erase the existence of its authors. It simply placed their work into a dataset where the books appeared as text first and books only after investigators worked backwards to identify them.
That is where my panic about AI scanning books really begins. The fear is not only ChatGPT being “basically high-tech plagiarism” as Noam Chomsky huffed earlier this year; that machines are reading and extracting what humans have written. It is that, at a certain scale, the culture of authorship itself starts to look like an inconvenient layer between a piece of information and the company that wants to acquire it. The strange orders for obsolete manuals, the 3,000-title lists, the 191,000 books in Books3 and the billions of agent actions all belong to the same expanding story. Human knowledge is being gathered at a scale that would have been unimaginable a few years ago, while the systems processing it are becoming harder even for their creators to monitor.
Perhaps the most revealing book in this whole story is therefore not a bestseller at all. It is, quite surprisingly, the unwanted manuscript that has sat for years in that second-hand shop we rarely ever step foot into. Maybe our job is to resist that logic while we still can, to keep asking whose words are being taken and what they become once they leave the page. This is a hard orientation towards asking and thinking about topics such as consent, compensation, and provenance. These are difficult topics, but if we stop asking these questions, the books will still be there; it is our idea of what a book is for that will have changed.