By now, it is widely understood that the large language models driving ChatGPT, Gemini, Claude, and similar chatbots are trained on vast repositories of published material—encompassing hundreds of millions of books, articles, academic papers, and essentially any publicly accessible internet content. The majority of published authors have unknowingly and without consent supplied the raw data for AI systems that now risk disrupting their professions. At first glance, this appears to be a clear violation of copyright law.
The legal reality, however, is far more nuanced.
“One of the challenges at the intersection of law and technology is the sheer volume of moving parts,” said Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, in an interview with TechCrunch. “The issues are highly complex, and strong sentiments exist on all sides.”
In a landmark ruling last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to a group of authors whose works were used to train the company’s models. While the decision initially appeared to favor writers, Judge Alsup determined that the AI training process itself was lawful. The penalty stemmed specifically from Anthropic’s acquisition of those books via illicit ‘shadow libraries.’
“Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them—but to turn a hard corner and create something different,” the judge wrote, likening the ingestion of trillions of words by an LLM to a writer’s study of literature.
Gellis views the ruling as largely beneficial to AI developers. For a company projecting roughly $200 billion in annual revenue by 2028, a $1.5 billion fine represents a manageable cost of business.
“The decision is generally positive for AI training because the judge analogized the process to reading a copyrighted work rather than copying it,” Gellis explained. “Copyright law centers on the act of copying; it does not prohibit the use, experience, consumption, or reading of a work.”
Copyright law has not been substantially updated since 1976, forcing judges to apply five-decade-old statutes to legal questions that could define the trajectory of the AI industry.
“There is widespread anxiety because the legal landscape is fragmented, and that fragmentation stems directly from this core question,” said Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International. “Stakeholders recognize that AI models have been trained on massive datasets, yet the law has not caught up to that reality.”
These disputes frequently center on fair use doctrine—specifically, whether the use of copyrighted material is sufficiently ‘transformative’ to be deemed legally permissible.
Fair use is a statutory exception permitting the unlicensed use of copyrighted works for purposes such as criticism, parody, education, and commentary. Courts evaluate fair use claims by weighing factors including the purpose and character of the use, the amount and substantiality of the portion used, and the effect on the potential market for the original work.
“Copyright fundamentally aims to protect and expand markets,” Henderson noted. “Court reasoning in AI cases varies significantly. A prevailing trend suggests that if a developer trains on proprietary data with the intent to directly compete with the source, courts view that unfavorably. Conversely, if the use does not pose a direct competitive threat, courts are more inclined to deem it acceptable.”
Henderson referenced the case of Thomson Reuters v. Ross Intelligence, in which the media and technology company sued the research firm for using its content to build a competing AI-driven legal platform.
“Ross’s use is not transformative because it lacks a ‘further purpose or different character’ than that of Thomson Reuters,” Judge Stephanos Bibas wrote in his opinion.
Judge Bibas ruled that training on Reuters’ content to create a direct competitor did not constitute fair use. Although authors might argue that chatbots compete with them by generating synthetic books from their works, that theory has not yet succeeded in court.
Regarding the intersection of AI and copyright, Gellis emphasizes the importance of distinguishing between two distinct issues: the copyright implications of training AI models versus the copyrightability of AI-generated output.
In Thaler v. Perlmutter, the court held that works generated entirely by AI are not eligible for copyright protection, raising difficult questions: how can one definitively determine whether a work was AI-generated, and if so, what proportion was created or assisted by AI?
“If you write a novel in Microsoft Word and use spell check, we comfortably accept that Word does not own your novel,” Gellis said. “AI is compelling us to confront a host of legal distinctions we have long deferred.”
Most major AI developers remain entangled in ongoing litigation over these issues, meaning definitive legal resolution remains distant.
“Early rulings are influential, but their impact could be reversed if other courts reach different conclusions; it will take later stages of litigation to establish prevailing precedent,” Gellis observed. “In the interim, these decisions are shaping the entire landscape. It would be unwise for AI companies to disregard them.”
Also Read
- Man City vs Bournemouth FREE Streams: How to watch Premier League 2026/27 from anywhere in the world
- Western Forests Face Permanent Conversion to Shrubland as Wildfires Intensify
- How Grasses Supercharged Their Starch and Biomass Production
- Spinning Stars and Fading Flares: New Insights into Black Hole Tidal Encounters

