Home/Technology/AIs can generate near-verbatim copies of novels from training data

AIs can generate near-verbatim copies of novels from training data

TechnologyFebruary 23, 20261 min readAttributed summary
AIs can generate near-verbatim copies of novels from training data
LLMs memorize more training data than previously thought.
Reading Settings

The world’s top AI models can be prompted to generate near-verbatim copies of bestselling novels, raising fresh questions about the industry’s claim that its systems do not store copyrighted works.

A series of recent studies has shown that large language models from OpenAI, Google, Meta, Anthropic, and xAI memorize far more of their training data than previously thought.

AI and legal experts told the FT this “memorization” ability could have serious ramifications on AI groups’ battle against dozens of copyright lawsuits around the world, as it undermines their core defense that LLMs “learn” from copyrighted works but do not store copies.

Read full article

Comments

Source: Ars Technica

Related technology stories

The Download: 10 climate tech companies to watch
Sourced report
technologyreview.com4 hours ago

The Download: 10 climate tech companies to watch

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. 10 climate tech companies to watch Each year, MIT

5 min briefingRead signal →