Regulation

AI-generated text

Uncertain Copyright Law Shapes AI Training Disputes

Courts are wrestling with whether using copyrighted works to train large AI models is lawful, with early rulings offering mixed signals.

Uncertain Copyright Law Shapes AI Training Disputes

Large language models that power tools like ChatGPT, Gemini, and Claude are trained on vast collections of published material — books, online articles, academic papers and other internet‑available content. Many authors have contributed to those datasets without being asked, raising the question whether that practice violates copyright and threatens creators’ livelihoods.

The answer is not straightforward. Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, told TechCrunch the area is complex and emotionally charged on both sides.

Early court rulings offer mixed signals

Last year Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers whose works were used in training the company’s AI models. On the surface that looked like a win for authors, but Alsup also found that Anthropic’s training was lawful — his penalty targeted the fact that Anthropic had obtained books from illegal online shadow libraries.

Alsup compared how a large language model ingests trillions of words to a writer studying literature, saying the models train to “turn a hard corner and create something different” rather than simply copy. Gellis views the ruling as relatively favorable for AI companies: she noted that a $1.5 billion fine may be manageable for a company that some project could reach roughly $200 billion in annual revenue by 2028.

Fair use and market impact are central

Much of the litigation turns on fair use — specifically whether the use of copyrighted works is sufficiently “transformative.” Fair use is a copyright exception that permits certain unlicensed uses for purposes such as criticism, parody, education and commentary. Courts weigh factors like purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the market.

Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch that courts have varied in their reasoning. When training or use directly competes with a copyright owner’s product, courts tend to find against fair use; when no direct competition is created, judges are more likely to accept the practice. Henderson pointed to the Thomson Reuters v. Ross Intelligence case as an example: Judge Stephanos Bibas concluded Ross’s use was not transformative because it served a purpose and character substantially similar to Thomson Reuters and produced a competing AI legal research platform.

Who owns what when AI is involved?

Another legal knot is whether AI‑generated or AI‑assisted works are eligible for copyright. In Thaler v. Perlmutter, a court ruled that an entirely AI‑generated work — one created 100% by AI — is not copyrightable. That raises difficult evidentiary questions: how can courts determine whether a work was produced with AI, and if so, to what degree?

Gellis drew an analogy to long‑accepted forms of assistance: people are comfortable saying that spell‑check in Microsoft Word doesn’t own your novel. But AI forces reconsideration of where the line is drawn between acceptable assistance and creative authorship.

Ongoing litigation, uncertain precedents

Most AI companies remain embroiled in pending litigation, so definitive legal answers are not imminent. Gellis observed that early decisions are influential but not final: subsequent courts could decide differently, and only later stages of litigation will reveal which legal theories prevail. In the meantime, those rulings shape industry behavior and cannot be ignored by AI developers.

Why this matters

The unsettled legal landscape affects both creators and AI builders. Authors worry about market displacement when AI systems use their works to generate new content; AI firms face legal and business risks depending on where source material came from and how courts classify that use. With U.S. copyright law substantively unchanged since 1976, judges are forced to interpret decades‑old rules in the context of technologies that those statutes did not contemplate, producing a patchwork of precedents that will take time to resolve.