IP & Rights

USA Today Sues OpenAI: What the $250M Copyright Case Means for Creators

2026-10-10 · 6 min read · AiDocX Newsroom

Seed story: "USA Today sues OpenAI for copyright infringement over AI training" (Reuters) · search original Written from facts verified across 2 news report(s) — original explainer, not a copy or translation. Sources listed at the end.

USA Today Co. has filed a lawsuit against OpenAI in the Southern District of New York, alleging that the company used over 160,000 of its copyrighted articles to train its GPT-2 models without authorization. As part of a wave of similar suits against major tech firms, this case seeks damages exceeding $250 million and could fundamentally reshape how creators and platforms navigate the legal boundaries of AI training data.

The Lawsuit: USA Today vs. OpenAI

On October 8, 2026, USA Today Co. and affiliated newspapers filed suit against OpenAI in the Southern District of New York. Represented by Steven Lieberman, the plaintiffs allege that OpenAI trained its large language models using hundreds of thousands of copyrighted articles without authorization. The complaint specifically cites internal data indicating that over 160,000 entries from their publications appear in the WebText corpus used for GPT-2 training.

The case, USA Today Co. v. OpenAI Foundation (No. 1:26-cv-08892), demands damages exceeding $250 million and an injunction to halt the alleged infringement. Key allegations include:

  • Microsoft providing OpenAI with a copy of the Bing Index.
  • Microsoft operating a crawler called Project Mango on OpenAI's behalf.
  • GPT-5.6 outputs allegedly replicating the plaintiffs' journalism.

For creators, this filing underscores the financial stakes of unauthorized AI training. By seeking substantial monetary relief, the lawsuit signals that publishers view AI scraping as a direct threat to their revenue streams and intellectual property rights.

Evidence of Systemic Scraping

The complaint moves beyond general allegations to present specific internal data, detailing the scale of the alleged infringement. According to the filing, content from USA Today and its affiliated newspapers comprises more than 160,000 entries within OpenAI’s WebText corpus. This dataset was reportedly used specifically for training the GPT-2 model, suggesting a deep integration of the plaintiffs' journalism into the foundational architecture of the AI.

This volume of data raises critical questions about how the material was acquired. The lawsuit alleges that Microsoft facilitated this collection by providing OpenAI with a copy of the Bing Index. Furthermore, it claims Microsoft operated a crawler known as Project Mango on OpenAI's behalf to gather this content.

  • 160,000+ entries from plaintiff publications found in the WebText corpus
  • Bing Index reportedly shared by Microsoft to aid data collection
  • Project Mango alleged to be a crawler operated for OpenAI

For creators, this evidence suggests that training data may not be gathered through isolated, manual efforts but via systematic, corporate-scale infrastructure. If these allegations hold, it implies that rights holders may have had no practical means to opt out of such automated, large-scale scraping.

The Core Legal Conflict

At the heart of USA Today Co. v. OpenAI is a fundamental clash over intellectual property boundaries. The plaintiffs argue that training large language models on copyrighted journalism constitutes direct infringement, rejecting the defense that such use qualifies as fair use. This dispute centers on whether the transformative nature of AI training justifies the unauthorized reproduction of protected works.

The complaint provides concrete evidence to support this legal stance. It highlights that the plaintiffs' content accounts for more than 160,000 entries in the WebText corpus used for GPT-2 training. Furthermore, the filing includes exhibits of copyright registrations alongside specific examples of GPT-5.6 outputs that allegedly replicate the plaintiffs' journalism.

  • Internal data showing extensive corpus inclusion
  • Alleged Microsoft assistance via Bing Index and Project Mango
  • Specific GPT-5.6 outputs mimicking original articles

For creators, this conflict is critical. If courts rule against fair use, it could fundamentally alter licensing agreements and payment structures, requiring explicit consent for any AI training on existing content.

A Wave of Litigation

The USA Today filing is part of a broader legal surge, with dozens of copyright owners now suing tech giants like OpenAI, Anthropic, and Meta. This wave of litigation signals a fundamental shift in how intellectual property is treated in the AI era, moving from isolated disputes to a systemic challenge against major platforms.

For creators, this trend highlights the fragility of current licensing models. As courts evaluate these cases, the legal landscape is rapidly evolving, potentially redefining what constitutes fair use for training data.

Key aspects of this shifting environment include:

  • Multiple high-profile suits targeting different AI developers simultaneously.
  • Allegations of unauthorized use of copyrighted material across various platforms.
  • A growing body of legal precedent that could influence future industry standards.

These developments suggest that the rights of content owners are becoming a central battleground in the AI industry, with potential long-term impacts on how creators are compensated and protected.

Implications for Creator Contracts

The USA Today lawsuit, seeking over $250 million in damages, signals a critical shift in how platforms must negotiate data rights. As courts evaluate whether unauthorized scraping constitutes infringement, existing licensing agreements may require urgent renegotiation. This legal pressure forces a re-evaluation of how independent professionals and media outlets define their intellectual property boundaries.

For creators, this means contracts must explicitly address AI training usage. Key considerations include:

  • Explicit Licensing Terms: Clearly defining whether content is licensed for model training.
  • Data Usage Rights: Specifying if data can be used for commercial AI products.
  • Compensation Structures: Establishing fair payment models for content inclusion.

According to reports, these clarifications are essential to protect revenue streams. Without precise contractual language, creators risk losing control over their work’s commercial value in the AI era.

Protecting Your Work in the AI Era

Audit Your Digital Footprint

As litigation like USA Today Co. v. OpenAI progresses, creators must proactively audit their digital presence to understand how their work circulates online. This involves identifying where your content is hosted and whether it is easily accessible to automated crawlers. Understanding the specific licensing terms of your platforms is equally critical, as standard agreements may inadvertently permit broad data usage.

To safeguard your intellectual property, consider these actionable steps:

  • Review your publishing contracts for clauses regarding AI training and data scraping.
  • Check if your content appears in public datasets or AI training corpora.
  • Consult legal counsel to interpret "fair use" boundaries in your specific context.

By taking these measures, you can better protect your rights and ensure that your creative output is not exploited without proper compensation or authorization.

FAQ

Why did USA Today sue OpenAI?

USA Today Co. and affiliated newspapers sued OpenAI for allegedly using hundreds of thousands of copyrighted articles to train its large language models without authorization. The complaint cites internal data showing that over 160,000 entries from the plaintiffs' publications were included in the WebText corpus used for GPT-2 training.

How much money is USA Today seeking in the lawsuit?

The plaintiffs are seeking damages exceeding $250 million from OpenAI. They are also requesting a court order to stop the alleged copyright infringement.

What role did Microsoft play according to the lawsuit?

The filing alleges that Microsoft provided OpenAI with a copy of the Bing Index to aid in data collection. Additionally, the complaint states that Microsoft operated a crawler called Project Mango on OpenAI's behalf.

Sources

Draft any contract in minutes — not billable hours

AiDocX generates artist, producer, influencer and crew agreements from a single prompt, then gets them e-signed. Free to start.

Try AiDocX free →

Related contract templates

← All briefings

Need a creator contract? AiDocX drafts & e-signs it in minutes. Try free →