Major artificial‑intelligence companies are confronting intensifying legal challenges over how they acquire and use creative works. The disputes span copyright, open‑source licensing, and antitrust concerns, painting a picture of a business model that relies on large‑scale data collection while resisting compensation to original creators.
A high‑profile case involves The New York Times suing OpenAI and Microsoft. According to unsealed court documents cited by 404 Media, Microsoft’s Director of Applied Science, Dr. Brent Hecht, described the lawsuit as “an astonishing theft of unprecedented proportions” and possibly “the largest theft of labor in human history.” The plaintiffs’ 92‑page brief quotes Hecht’s remarks, and an internal Microsoft policy document acknowledges that generative AI “could significantly disrupt the employment of the very people who generated the data on which the foundation model was trained” because LLMs “are a product that destroys its supply chain.” This admission highlights the risk of “model collapse,” where AI erodes the very content ecosystem it depends on.
Judge Sidney H. Stein of the U.S. District Court for the Southern District of New York has yet to rule. OpenAI and Microsoft argue that training on publicly available articles and books constitutes fair use, advancing public knowledge without unlawfully substituting existing markets. The plaintiffs counter that the defendants knowingly harvested copyrighted material without regard for licensing or remuneration.
Depositions from OpenAI’s corporate representative reveal a striking lack of safeguards. The witness testified that the company was unaware of any effort to detect or exclude paywalled content during model training. When OpenAI co‑founder Greg Brockman was informed that the system could bypass paywalls, his response was recorded as “ah, nice.”
Beyond journalism, the legal battle touches on software development. In *Doe v. GitHub*, the Ninth Circuit issued a narrow victory for GitHub, Microsoft, and OpenAI, holding that the plaintiffs failed to prove a violation of a specific DMCA provision concerning the removal or alteration of copyright‑management information (CMI). Judge Eric Miller wrote that creating new code without CMI does not constitute “removal” or “alteration.” However, the ruling does not settle broader questions about fair‑use training on open‑source projects, the applicability of licenses such as GPL or Apache, or whether generated code can infringe copyright.
An open‑source attorney warns that many cases are beginning to treat licenses as contracts rather than as intellectual‑property grants. This shift could undermine the foundations of open‑source licensing, which traditionally rely on ownership and infringement analyses. If licenses are enforced merely as contractual promises, the legal framework that has long supported collaborative software development may be destabilized.
The fallout extends to antitrust. A federal suit alleges that Anthropic, OpenAI, SpaceXAI, and Google colluded to “pace” frontier AI development, effectively limiting competition. Executives—including Anthropic’s Dario Amodei, OpenAI’s Sam Altman, Google DeepMind’s Demis Hassabis, and Elon Musk—have publicly advocated for coordinated slowdowns on safety grounds. While safety is a legitimate concern, the complaint argues that such coordination amounts to competitors dictating market terms, which could stifle innovation and entrench dominant players.
Antitrust law requires proof of an agreement and tangible harm to competition. Public statements alone are insufficient, and even successful suits have historically struggled to impose meaningful penalties, as seen with Google’s ad business and prior search‑related settlements.
These overlapping legal fronts reveal a pattern: AI giants routinely invoke fair‑use defenses for copyrighted content, claim generated works are distinct from source material, and cite safety coordination to justify market control. The cumulative effect is a concentration of data, market share, and revenue among a few powerful firms, while creators and developers risk losing income, attribution, and legal recourse.
The core issue is not whether AI will advance, but whether its builders can treat others’ intellectual property as a free raw material to be harvested and then used to capture creators’ audiences, revenue, and bargaining power. Without decisive action to curb this data grab, the benefits of AI may be confined to a tiny elite, leaving the broader populace to navigate a new form of digital feudalism.
<img src="https://image.theregister.com/5299100.webp?imageId=5299100&x=0.00&y=0.00&cropw=100.00&croph=100.00&width=960&height=644&format=jpg" width="480" height="322" alt="Colorful graffiti mural beside a tree-lined street with parked cars in Warsaw, Poland." loading="lazy" style=""/>
<p>
<figcaption itemprop="caption" class="">Warsaw, Poland: A stylized interpretation of the ancient Egyptian god Thoth (also the self-styled tag of the artist), states: "Nie bój się niczego, bowiem wszystko jest twoim," translating to: "Fear nothing, for everything is yours."</figcaption>
<figcaption itemprop="author" class="" data-byline-prefix="">Pic credit: Anton Kustsinski/Shutterstock</figcaption>
</p>
The emerging legal landscape underscores the urgent need for clear rules governing AI training data, licensing compliance, and competitive practices. Until those safeguards are established, the current trajectory risks entrenching a system where a handful of trillionaires reap the rewards while the rest become digital serfs.
Also Read
- Record‑Breaking Approaches: The Nearest Object to the Sun
- Berlin Marathon 2026: Where to Watch the World’s Fastest Marathon Live Online
- Elite Journal Acceptance Rates Skewed by Institutional Prestige, Research Topic, and Team Size, Large-Scale Analysis Shows
- PNOĒ Launches Self‑Serve Breath‑Analysis Mask for Lab‑Grade Metabolic Testing


