Justified Posteriors
Justified Posteriors
Should AI Read Without Permission?
0:00
-55:05

Should AI Read Without Permission?

Discovering what AI learns from unlicensed books in "Cloze Encounters: The Impact of Pirated Data Access on LLM Performance" by Stella Jia and Abhishek Nagaraj

Many of today’s thinkers and journalists worry that AI models are eating their lunch: hoovering up these authors’ best ideas and giving them away for free or nearly free. Beyond fairness, there is a worry that these authors will stop producing valuable content if they can’t be compensated for their work. On the other hand, making lots of data freely accessible makes AI models better, potentially increasing the utility of everyone using them. Lawsuits are working their way through the courts as we speak of AI with property rights.

Society needs a better of understanding the harms and benefits of different AI property rights regimes.

A useful first question is “How much is the AI actually remembering about specific books it is illicitly reading?” To find out, co-hosts Seth and Andrey read “Cloze Encounters: The Impact of Pirated Data Access on LLM Performance”. The paper cleverly measures this through how often the AI can recall proper names from the dubiously legal “Book3” darkweb data repository — although Andrey raises some experimental concerns.

Listen in to hear more about what our AI models are learning from naughty books, and how Seth and Andrey think that should inform AI property rights moving forward.

Also mentioned in the podcast are:

Discussion about this episode

User's avatar

Ready for more?