'Unlicensed, unrestricted AI training could destroy the ecosystem for books' — quote of the day by the Authors Guild on the sourcing of training data
Many of today's widely used AI systems have been trained on material sourced from books
To achieve any level of competency, large language models (LLMs) need ample data for sufficient training. AI companies have looked to various sources to mine this information, including content publicly available on the internet, synthetic data generated from other AI models, and printed literature.
Reading difficulties
Prompted by news that AI companies were allegedly using books from pirate ebook sites to build their LLMs, writers, authors, and publishers publicly called out this deeply worrying process.
This article is part of TechRadar Pro's QOTD project to provide an insight into the minds of the brightest and most recognized figures in the technology industry today and in years gone by. Read the full series here.
The professional organization known as the Authors Guild responded to various stories about AI companies scanning books to train their AI models (both illegally and legally) with incredibly comprehensive guidelines on AI licensing.
This document covered the various manifestations of the use of published works by AI companies, including its legal perspective on the legitimacy of using such works. It also highlighted that the continued data harvesting processes would risk destroying the ecosystem for books that currently exists.
Book buying
The use of books by AI companies is an ongoing concern. But in recent months the focus has pivoted to those that buy, scan – and destroy – books on an industrial scale.
For example, court documents revealed the existence of 'Project Panama' inside Anthropic. This is a scheme in which the company aims to "destructively scan all the books in the world" and used a codename because "we don’t want it to be known that we are working on this.".
To train Claude, Anthropic had to procure a large and high-quality dataset, so it set out to purchase books on an industrial scale because of the relatively high-quality nature of the writing compared with, say, writing found online.
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

Keumars Afifi-Sabet is a freelance contributor for Tech Radar and the Technology Editor for Live Science. He has written for a variety of publications including ITPro, The Week Digital and ComputerActive. He has worked as a technology journalist for more than five years, having previously held the role of features editor with ITPro. In his previous role, he oversaw the commissioning and publishing of long form in areas including AI, cyber security, cloud computing and digital transformation.
You must confirm your public display name before commenting
Please logout and then login again, you will then be prompted to enter your display name.