Australian dealers say AI book scanning destroys rarities

Australian dealers say AI book scanning destroys rarities

On August 1, 2026, The Guardian reported that Australian secondhand booksellers fear rare titles are being scanned and then destroyed to feed AI models. The sellers believe they were pulled into an unregulated data pipeline, where old books are bought in bulk, digitized, and shredded after capture. The practice — call it AI book scanning — shifts the AI data fight from the web to the stacks, with real-world losses that can’t be undone.

What The Guardian uncovered about AI book scanning

According to The Guardian’s AI section on August 1, 2026, dealers in Australia raised the alarm after noticing patterns that suggest buyers are funneling secondhand and rare books into industrial scanning operations. The claim is stark: titles with cultural and scholarly value aren’t just being copied — they’re being destroyed once the pages are captured for data. That allegation reframes the AI training data debate. It is no longer only about what’s scraped online or covered by licenses. It’s about whether physical artifacts are being sacrificed to feed models.

There’s logic to the reported practice. Scanning unlocks high-quality, long-form text with clean pagination and front matter that can be mapped to metadata. Shredding lowers storage costs and avoids resale, which could trace the item back to a buyer. None of that makes it acceptable. It just makes it efficient. And it keeps the softest spot in the chain — the physical book — out of sight.

Why shredding scanned books threatens cultural memory

Destroying a book after digitization isn’t a neutral act. It erases bindings, marginalia, inscriptions, paper stock, and publishing quirks that carry meaning. Those signals matter to researchers, librarians, and historians. UNESCO’s Memory of the World program stresses preservation of documentary heritage and long-term access. Mass scanning that ends in a shredder cuts against that goal, especially for rare or out-of-print works that may have few surviving copies.

Libraries have built norms around this. They digitize to widen access, then keep or responsibly deaccession the original using clear policies and records. Australia has done celebrated, open digitization at scale through Trove at the National Library of Australia, pairing access with stewardship. What The Guardian describes is something else: private, opaque acquisition where the output is data for AI systems, and the physical source disappears. That secrecy invites bad incentives and accidental harm.

The public also loses a chance to test claims about provenance and consent. Without a record of what was scanned, when, and under what terms, there’s no way to separate lawful digitization from material obtained through aggressive or deceptive buying. In practice, that leaves honest booksellers exposed and collectors wary.

The legal gray zone around training data

Publishers and authors have warned for years that AI training on books raises legal and moral questions. In 2023 and 2024, authors filed lawsuits against AI companies over the use of books in training sets, arguing unauthorized copying and derivative use. The Authors Guild summarizes several of those cases on its site. Courts are still hashing out where fair use ends and infringement begins in the training context.

A different but relevant fight played out in Hachette v. Internet Archive, where a U.S. court rejected a library’s broader lending theory for scanned books. That case wasn’t about model training, but it shows how judges scrutinize large-scale book scanning without clear licenses. If AI firms or their intermediaries are sourcing physical copies in bulk, scanning them, and then pulping the originals, the chain of custody and consent questions don’t just persist. They multiply.

Australia’s policy community is watching, too. The Attorney-General’s Department has been reviewing copyright and AI, including how training data is obtained and documented; see its consultation materials on copyright and artificial intelligence. None of that directly settles The Guardian’s reported scenario. It does underline the need for traceability and standards, or at least for companies to prove they didn’t rely on destructive practices masked as routine procurement. Put bluntly, AI book scanning that ends in a shredder is a compliance and reputational hazard, not just a PR problem.

What would curb the damage from scanning books for AI

Start with chain of custody. AI developers should require their data brokers and contractors to certify how texts were obtained, including whether physical originals were retained or responsibly transferred. Procurement codes can embed a simple rule: no datasets built from destroyed originals unless a public institution or rights holder authorized the process and kept a preservation copy.

Next, adopt “scan-and-return” as a default. Libraries have done it for decades. If a private buyer needs a text for research, scan it with the seller’s consent, then return it in the same condition. For fragile or rare works, involve conservation professionals. That won’t satisfy every AI training need, but it draws a bright line that deters the worst behavior.

Dealers can help, too. They can ask bulk buyers for a use declaration, track batches, and flag unusual requests to trade associations. A basic form that captures identity, intended use, and resale terms would add friction without smothering legitimate trade. It also gives honest buyers a way to distinguish themselves from data scavengers.

Regulators don’t need to write a whole new rulebook to move the market. Existing consumer law and unfair practices rules already cover deception in procurement. Culture agencies can issue guidance that books sold by public institutions, charities, and estate vendors shouldn’t be destroyed as part of private digitization without oversight. Grant programs can fund better dealer training and simple provenance tools. If the state funds digitization, it can require public logs of what gets scanned and how.

Finally, AI firms should publish training data intake policies with one line readers can check: we don’t source from destructive pipelines. Independent auditors can test that claim by sampling datasets for telltale metadata and verifying supplier practices. The same companies that build safety benchmarks can help the field agree on what a clean, documented text dataset looks like.

What this means for the next phase of AI

The story The Guardian surfaced isn’t a curiosity. It signals the end of the “free data” era and the start of hard choices about where the next tokens come from. Web text is messy, licensed corpora are pricey, and long-form, rights-cleared writing is scarce. That pressure pushes some actors to cut corners. When the corner is a physical book, the damage is visible — and irreversible.

If the industry wants public trust, it has to prove that its hunger for text isn’t chewing through cultural heritage. That starts with visible standards, credible audits, and a willingness to say no to gray-market supply. The fastest way to show good faith would be for major labs to renounce AI book scanning from destructive sources and to back dealer-led safeguards. Failing that, the scandal moves from the stacks to the courtroom, and the harm will already be done. For more on this, see bloomberg.com and nytimes.com.

Related reading: AI in EducationData PrivacyAI in Society