On August 15, 2026, The Guardian reported that secondhand booksellers in the UK and Ireland suspect AI firms are behind “strange” bulk orders, following revelations that Anthropic had spent millions on books for data acquisition. The surge points to a new, opaque market: AI training data books sourced from the resale trade rather than from publishers.
Why AI training data books are spiking
The suspicion from shop owners, reported by The Guardian, slots into a broader scramble for long-form text to train models. Web data is abundant, but licenses are contested and opt-out signals are spreading. Books offer dense, edited prose across topics and eras. Out-of-print titles add variety that a web crawl can miss. Put together, the incentive to buy quietly and at scale is clear.
There’s also a timing tell. As lawsuits and licensing talks drag on, buying used books can look like a faster route to assemble training corpora. US authors have already sued AI companies over book use, a trend documented by the Authors Guild. According to The Guardian’s summary, Anthropic’s spending in the millions signals that acquiring physical copies, then scanning, is no fringe tactic—it’s a budget line.
What bulk buying means for the secondhand trade
Shops built to serve readers and researchers are being pulled into a data supply chain they don’t see. Orders that look like normal wholesale bundles can mask a single end use: ingestion into a training set. That can push up prices for sought-after fields, drain shelves of certain academic or technical titles, and leave local buyers empty-handed. If the buyer refuses transparency, the shop carries the reputational risk with its community.
Small businesses also bear operational exposure. High-volume, one-off orders raise fraud and chargeback risk. Buyers using intermediaries make returns harder to manage. And if labs drive heavy demand for particular ISBNs, pricing signals get distorted, then snap back once the project ends. In other words, AI training data books don’t just move quietly through warehouses; they reshape local markets while they do.
The legal gray zone around scanning and mining
Copyright lines matter here. In the UK, the text and data mining exception covers non-commercial research. Commercial mining still requires permission from rights holders. Government guidance spells this out in the text and data mining copyright exception. Scanning a book makes a copy, which engages copyright; using the text to train a model can be a separate legal question. If labs rely on used copies to scan at scale, they still face the permission hurdle for commercial mining.
That’s why the purchasing pattern matters. If an AI company is sourcing hundreds of volumes from the secondhand channel, it may be avoiding explicit licensing pathways with publishers or authors. According to The Guardian’s account, the link between bulk buys and later scanning is already part of the discussion around Anthropic. Even where fair dealing, exceptions, or licenses apply, the chain of custody should be clear. Today, it often isn’t.
How sellers can respond without losing the sale
Independent shops don’t need to become investigators. They can tighten the basics: require verifiable identities for large orders, set volume caps on rare or specialist titles, and keep a watch list of ISBNs that see sudden spikes. Where buyers are open, sellers can ask about intended use, then steer them to publisher channels for licenses when the answers point to commercial text mining.
Trade bodies could help standardize that playbook. A simple disclosure clause for bulk orders—declaring whether the purchase is for reproduction or machine ingestion—would give shops a choice about supply and pricing. Downstream, provenance signals will matter too. If labs tout safety and transparency, they can document sources and consent, much like the C2PA standard does for content credentials. The same logic applies to training data.
Why this isn’t just a publishing story
The Guardian’s report captures a pivot point. AI model builders don’t only need GPUs and power; they also need lawful, diverse text. That makes AI training data books a commodity class of their own, tugging on the secondhand market and on copyright norms. The surprise for many will be where the pressure shows up first: at the till in a local shop, when a stranger asks for ten copies of the same niche title.
If this becomes standard practice, expect clearer guardrails. Regulators will push for transparency in data sourcing. Sellers will bake in buyer verification. And labs with scale will weigh the costs of quiet bulk buying against the stability of licensed pipelines. Either way, AI training data books have entered the open market, and the market is noticing. For more on this, see anthropic.com and bloomberg.com.
Related reading: Video Generation • AI Agents • AI Tools & Platforms
