When an AI model trained on millions of copyrighted works generates a novel, a painting, or a piece of code — who owns the output? The creator of the training data? The company that built the model? The user who wrote the prompt? Or nobody? This question is being litigated in courts worldwide, and the outcomes will define the economics of the entire generative AI industry.
The Key Lawsuits
The New York Times v. OpenAI and Microsoft (filed Dec 2023): The most consequential AI copyright case. The Times alleges that ChatGPT and Bing Chat can reproduce near-verbatim passages from its articles, constituting copyright infringement on a massive scale. OpenAI’s defense: AI training on publicly available text constitutes “fair use” — a transformative use that doesn’t substitute for the original. The case is proceeding through the Southern District of New York and a ruling could come in 2026. If the Times prevails, every AI company training on web-scraped data faces retroactive liability.
Getty Images v. Stability AI (UK and US): Getty alleges that Stable Diffusion was trained on 12 million Getty images without license. Early Stable Diffusion outputs sometimes included distorted Getty watermarks — evidence that the training data was copied rather than merely “learned from.” Authors Guild v. OpenAI: A class action representing thousands of authors whose books were allegedly used to train GPT models without compensation. Sarah Silverman, Christopher Golden et al. v. Meta: Alleging that Meta’s LLaMA was trained on pirated book datasets.
The Fair Use Question
US copyright law’s “fair use” doctrine considers four factors: the purpose and character of the use (is it transformative?), the nature of the copyrighted work, the amount used relative to the whole, and the effect on the market for the original. AI companies argue: training is transformative because the model doesn’t store or reproduce works — it learns patterns. Critics argue: the effect on the market is massive — AI-generated articles directly compete with the original journalism, art, and writing that trained them. The courts will likely rule differently depending on whether the AI output competes with the original. An AI that generates news articles in the style of the NYT competes directly; an AI that uses NYT-trained language understanding to power a customer service chatbot doesn’t.
The Emerging Licensing Landscape
While courts deliberate, the market is building licensing solutions. OpenAI has signed licensing deals with the Associated Press, Axel Springer, Le Monde, and several other publishers. Google has similar agreements. These deals — typically $5-50 million per publisher per year — establish a commercial model for AI training data. For smaller creators, platforms like Spawning.ai allow artists to opt their work in or out of AI training datasets. The likely equilibrium: large-scale training data will require licensing, but the specific terms and pricing will vary by content type, with premium journalism and professional photography commanding higher rates than user-generated content.
For more on AI’s legal and regulatory landscape, explore our AI Policy & Regulation coverage. For the broader AI industry landscape, browse our AI Tools & Applications section.
Further Reading
Related: Building a Waitlist That Converts: Pre-Launch Growth Strategies — Startup Nerve
Related: Growth Hacking 2026: AI Tactics, Viral Loops, Community — Startup Nerve
What to Watch Next
The trajectory of copyright wars owns content will be shaped by several converging factors over the next 12-18 months. Regulatory frameworks are catching up with technological capabilities, enterprise adoption is accelerating beyond pilot programs, and the talent pool is deepening as more engineers and researchers enter the field. The companies and research groups that establish technical leadership now will benefit from compounding advantages — better data flywheels, stronger talent networks, and deeper domain expertise. For industry observers and technology leaders, tracking the leading indicators discussed above will provide the clearest signals about which applications cross the threshold from promising to transformative.
Dive deeper: This article is part of our comprehensive guide — The State of AI in 2026: Everything You Need to Know.
