How AI book scanning is driving a boom in used-book sales

How AI book scanning is driving a boom in used-book sales

On August 15, 2026, BBC Technology highlighted a sharp rise in secondhand book sales across the UK and Ireland. That surge aligns with fresh reporting in The Guardian the same day that independent sellers suspect artificial intelligence firms are behind unusual bulk purchases. Put together, the two accounts point to a simple driver with big consequences: AI book scanning is becoming a quiet force in the used-book market.

Where AI book scanning shows up in the market

The BBC front page for technology on August 15, 2026 drew attention to booming secondhand sales, raising the question of whether AI is the cause (BBC Technology). That hypothesis got support hours later when The Guardian reported that booksellers in the UK and Ireland are fielding “strange” bulk orders and suspect AI companies are the buyers. The report followed prior revelations that Anthropic had spent millions of dollars on books to scan for “data acquisition,” a phrase that explains both the motive and the method.

For shop owners, the signals are hard to miss: large, targeted baskets that don’t look like a school’s term order or a collector’s haul, and that repeat across different sellers. While neither outlet named specific buyers on August 15, the pattern The Guardian describes matches the incentives of model builders racing to add long-form text to training sets.

The AI training data race and the used-book boom

Why would AI labs buy old paper instead of licensing e-books? Cost, coverage, and control. Secondhand copies are cheap relative to negotiating title-by-title rights. They can also fill gaps where digital editions are scarce, especially for backlist non-fiction, technical manuals, and out-of-print works. And once a company holds a physical copy, it can scan on its own schedule and quality standard, feeding a pipeline tuned for large-scale ingestion of long-form text.

There is another reason: legal uncertainty. The rules around text and data mining vary by jurisdiction and by use. In the UK, debates continue over how copyright exceptions apply to commercial AI training, with the government’s Intellectual Property Office hosting guidance and consultations that remain closely watched (UK IPO). Buying physical books does not make copyright questions vanish, but it can reduce contracting friction and keep procurement activity quieter than a marquee licensing deal.

That context helps explain why the BBC’s on-the-ground signal of brisk secondhand sales and The Guardian’s account of odd bulk purchases are two sides of the same coin. The demand source is new, and it scales.

Short-term winners, longer-term risks

Independent sellers look like short-term winners. Extra orders mean faster stock turnover and healthier margins, especially on backlist titles that used to sit for months. Niche subjects could see price spikes as supply tightens, which rewards well-curated shops but frustrates students and researchers who relied on low-cost used copies.

Authors are not obvious beneficiaries. If AI buyers are scanning without explicit licenses from rightsholders, creators may see none of the upside even as their works shape model behavior. The Guardian’s reference to Anthropic’s prior multimillion-dollar spend signals that major labs are willing to pay—just not always to authors or publishers. That mismatch will fuel calls for clearer licensing channels and auditing of training sets.

AI firms face reputational risk and possible legal exposure. Training on scanned books without firm legal footing invites challenges that can delay releases or force costly retraining. It also raises trust issues: enterprises and public bodies increasingly ask vendors to document how models were built and what went in, which puts opaque book-scanning programs under a brighter light.

What shops, publishers, and regulators will watch next

Booksellers may tweak policies to keep regular customers happy while meeting demand. Expect caps on single-buyer quantities for sensitive categories, closer review of repeated large orders, and coordination across marketplaces to spot patterns. Online platforms that host used-book listings could introduce basic anomaly checks to flag mass purchases that drain supply.

Publishers have a decision to make. They can sit back and let secondary markets feed AI, or they can package backlist rights into clear, priced bundles designed for training. That path is administrative work, but it converts uncertainty into revenue and gives them leverage over how content is used.

Regulators will keep probing the gap between what copyright law says and how AI companies behave in practice. In the UK, the IPO’s process suggests more guidance is likely, and parliamentary scrutiny could sharpen if retail buyers and university libraries start reporting shortages tied to bulk orders. Clearer disclosure rules—who is buying, in what volume, and for what use—would also reduce suspicion and stabilize prices.

For now, the signals line up. The BBC points to a real-world sales spike; The Guardian surfaces the likely buyer profile and ties it to earlier spending by a leading lab. That makes this more than a curiosity. It is an early market proof that model builders will pay for long-form text at scale—and that AI book scanning is already reshaping shelves, prices, and expectations. For more on this, see anthropic.com.

Related reading: NVIDIAMeta AIAI & Big Tech