The practice of AI companies destroying old books to train their models has sparked intense debate and raised ethical concerns. While the industry's hunger for diverse data sources is understandable, the methods employed are causing alarm among book enthusiasts and copyright holders alike.
One company, ISBNdb, initially offered a service that involved physically destroying millions of books, a process that raised eyebrows and sparked a heated discussion. The company's attempt to frame the destruction as a recycling initiative was met with skepticism, as the optics of burning libraries were hard to ignore. This incident highlights the delicate balance between innovation and preserving cultural heritage.
The AI industry's reliance on book data is not new, but the methods used to acquire it are. Large Language Models (LLMs) require vast amounts of text data, and initially, this was sourced from legally available online books or public domain works. However, the competition between AI companies like OpenAI and Anthropic has led to a darker approach.
Anthropic, in particular, was found to have used millions of pirated books to train its Claude model. This was revealed in a class-action lawsuit by the Authors Guild, which resulted in a record $1.5 billion copyright settlement. The lawsuit uncovered Anthropic's Project Panama, where they purchased and physically destroyed books, a process that was deemed 'transformative' and fell under fair use.
The scale of this operation was staggering, with Anthropic aiming to convert 500,000 to 2 million books in just six months. This raises questions about the efficiency and morality of such methods. While AI companies argue for the need to train their models, the destruction of physical books can be seen as a violation of cultural and intellectual property rights.
In contrast, organizations like the Internet Archive have adopted nondestructive scanning methods. They use human operators to turn pages, ensuring the preservation of rare books. This approach, while slower, maintains the integrity of the original text and avoids the ethical dilemmas associated with book destruction.
The debate surrounding AI companies' book destruction practices is complex. While the industry's rapid growth demands diverse training data, the methods employed can have severe consequences. The spotlight of public attention may be the only factor preventing further destruction, as seen with ISBNdb's pivot away from destructive scanning. This incident serves as a reminder of the importance of ethical considerations in the pursuit of technological advancement.