The Dark Side of LLM Training: How Anthropic's 'Project Panama' Reportedly Destroyed Millions of Books
Discover the disturbing details of Anthropic's 'Project Panama,' a reported effort to destructively scan millions of physical books to train Claude AI.

A Literary Massacre for Machine Learning
In the relentless pursuit of artificial intelligence supremacy, the hunger for high-quality training data has reached a fever pitch. While most users interact with AI through clean interfaces and helpful responses, a disturbing trend has emerged behind the scenes. Reports have surfaced indicating that Anthropic, the AI powerhouse behind the Claude series, engaged in a massive operation to acquire and systematically dismantle millions of physical books to accelerate the training of its large language models (LLMs).
The operation, allegedly codenamed "Project Panama," highlights a growing conflict between the preservation of human culture and the industrialization of data acquisition for the AI era.
Unveiling 'Project Panama': The Mechanics of Destruction
According to reports, including a detailed investigation by the Washington Post, Project Panama was not merely a digitization effort but a destructive one. The process was designed for maximum efficiency and speed, prioritizing throughput over preservation.
The workflow reportedly involved several ruthless steps:
- Bulk Acquisition: Spending millions of dollars to acquire vast quantities of printed literature from commercial markets.
- Destructive Preparation: Rather than carefully scanning bound volumes, the company allegedly cut off the spines of the books. This allowed individual pages to be fed through high-speed industrial scanners at a rate far exceeding traditional book-scanning methods.
- Digitization: The loose pages were converted into machine-readable text, which was then integrated into the massive datasets used to refine Claude's cognitive abilities.
- Disposal: Once the text was extracted, the remaining physical fragments—now stripped of their bindings—were discarded as waste.
An internal planning document allegedly described the project with chilling clarity: "Project Panama is our effort to destructively scan all the books in the world," adding, "We don’t want it to be known that we are working on this."
The Legal Fallout and the 'Fair Use' Debate
The controversy surrounding Project Panama is inextricably linked to the broader legal battle over copyright in the age of AI. Anthropic recently settled a massive copyright lawsuit for $1.5 billion. While the court ruled that training AI models generally falls under "fair use"—comparing the AI's learning process to a human writer studying existing works to create something new—the court did find that Anthropic violated copyrights by maintaining a library of 7 million pirated books.
While other AI giants like OpenAI, Google, and Meta have faced similar lawsuits regarding the use of digital libraries, Anthropic is currently the only major player reported to have physically destroyed physical copies of books to achieve its data goals.
Anthropic's Response
When questioned by Tom's Guide, a spokesperson for Anthropic defended the practice of sourcing books, noting it is a common approach across the AI industry. The company emphasized that they do not target rare or antiquarian volumes, stating, "None of our data acquisition programs buy and destroy ‘rare’ or ‘antiquarian’ books. We buy books from regular commercial markets." Regarding the legal disputes, the company pointed to the Bartz settlement, maintaining that the courts have upheld the legality of AI training under copyright law.
The Ethical Bottom Line
The revelation of Project Panama raises profound ethical questions. While the company claims they only targeted common commercial books, the systematic destruction of millions of physical objects for a digital copy creates a precarious precedent. The loss of physical literature, even in bulk, represents a shift toward a world where the original artifacts of human thought are sacrificed to fuel the algorithms that may one day replace the authors themselves.
As AI continues to evolve, the tension between the need for data and the respect for physical and intellectual property remains one of the most critical challenges for the technology sector.