All NewsEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
Story archiveAll categories
← All Stories

AI Companies Destroying Millions of Books for Training Data

Created at 29 Jul · 4:46 PM1 source↑ Market-relevant
IN SHORT

AI companies are anonymously purchasing physical books in bulk, destroying them after scanning for training data. This practice has surged demand for obscure titles, raising concerns about the permanent loss of rare books, despite recent fair use rulings.

✉Newsletter

PiQ Daily

Pick your topics. Get only what matters, on your cadence.

Key Numbers

$1.5 billioncopyright settlement by Anthropic
$3,000per book payment in Anthropic settlement

Who's Involved

AI companies
anonymously acquiring and destroying books for training data
Booksellers
reporting surged demand for obscure titles
Federal judge
ruling destructive scanning can be fair use
Anthropic
settling copyright lawsuit for $1.5 billion
Elon Musk
speaking out against destructive book scanning
AI Companies Destroying Millions of Books for Training Data

↳ Why This Matters

This practice raises significant ethical and legal questions about copyright, the preservation of human knowledge, and the potential for AI development to permanently erase cultural artifacts.

Key facts

  • AI companies are purchasing millions of physical books in bulk and destroying them after scanning for training data.
  • This practice has led to a surge in demand for obscure and out-of-print books.
  • A federal judge ruled that destructive scanning of legally purchased books can be considered fair use.
  • Anthropic will pay authors approximately $3,000 per book in a $1.5 billion copyright settlement.
  • Concerns exist about the permanent loss of rare and uncommon books due to this practice.

AI companies are engaging in a practice akin to book burning, anonymously purchasing millions of physical books in bulk, destroying them after scanning their content for AI training data. This surge in demand has significantly boosted sales for booksellers, particularly for obscure and out-of-print titles, leading to fears that rare books are being permanently lost.

A federal judge recently ruled that the destructive scanning of legally purchased books can qualify as transformative fair use, a decision that has implications for ongoing copyright litigation involving major AI firms like OpenAI and Meta. However, in a separate development, Anthropic has agreed to a $1.5 billion copyright settlement, paying thousands of authors approximately $3,000 per book for using pirated copies to train its AI model, Claude.

The practice has drawn criticism, with some viewing it as a race to preserve human-authored knowledge before it is overshadowed by AI-generated text. Prominent figures like Elon Musk have publicly opposed the destructive scanning method, advocating for more traditional scanning techniques to preserve rare books.

Frequently asked questions

AI companies are destroying books after scanning them to use the content as training data for developing more powerful AI models.

Demand for obscure and out-of-print titles has surged, leading to increased sales for booksellers but also concerns about the permanent loss of these rare books.

A federal judge ruled that destructive scanning of legally purchased books can qualify as fair use, though separate copyright litigation continues.

Anthropic agreed to a $1.5 billion copyright settlement, paying authors about $3,000 per book for using pirated works to train its AI.

What Happens Next

01Further legal rulings on copyright and fair use in AI training are expected.
02AI companies may face increased scrutiny and public pressure regarding their data acquisition methods.

Get the newsletter.

Pick the topics you actually care about. We'll email when there's news worth your time, on the cadence you choose. Cancel any time from your account.

Cadence

How It Developed

AI companies are anonymously acquiring physical books in bulk through intermediaries.
The acquired books are destroyed after being scanned for AI training datasets.
Demand for obscure and out-of-print titles has surged, impacting the used-book market.
A federal judge ruled that destructive scanning of legally purchased books can qualify as fair use.
Anthropic agreed to a $1.5 billion copyright settlement for using pirated works to train its AI.
Some AI developers, including Elon Musk, have spoken out against the destructive scanning practice.

Sources

T1
AI Book Burning? Companies Are Destroying Millions of Books to Feed ChatbotsDecrypt

Related Stories

AI firms target rare books for training data, sparking 'dystopian' concerns
29 Jul · 3:26 PM
AI Boom Fuels Demand for Electricians, Carpenters
29 Jul · 9:11 AM
PwC reports found to contain AI hallucinations and fabricated claims
29 Jul · 9:06 AM
Ofgem proposes fees for AI data centres to ease UK grid connection queues
29 Jul · 7:26 AM
Atlassian caps AI spending at $2,000/month per employee amid rising costs
29 Jul · 3:11 PM