USA Today Sues OpenAI: What the $250M Copyright Case Means for Creators
Seed story: "USA Today sues OpenAI for copyright infringement over AI training" (Reuters) · search original Written from facts verified across 2 news report(s) — original explainer, not a copy or translation. Sources listed at the end.
USA Today Co. has filed a lawsuit against OpenAI in the Southern District of New York, alleging that the company used over 160,000 of its copyrighted articles to train its GPT-2 models without authorization. As part of a wave of similar suits against major tech firms, this case seeks damages exceeding $250 million and could fundamentally reshape how creators and platforms navigate the legal boundaries of AI training data.
The Lawsuit: USA Today vs. OpenAI
On October 8, 2026, USA Today Co. and affiliated newspapers filed suit against OpenAI in the Southern District of New York. Represented by Steven Lieberman, the plaintiffs allege that OpenAI trained its large language models using hundreds of thousands of copyrighted articles without authorization. The complaint specifically cites internal data indicating that over 160,000 entries from their publications appear in the WebText corpus used for GPT-2 training.
The case, USA Today Co. v. OpenAI Foundation (No. 1:26-cv-08892), demands damages exceeding $250 million and an injunction to halt the alleged infringement. Key allegations include:
- Microsoft providing OpenAI with a copy of the Bing Index.
- Microsoft operating a crawler called Project Mango on OpenAI's behalf.
- GPT-5.6 outputs allegedly replicating the plaintiffs' journalism.
For creators, this filing underscores the financial stakes of unauthorized AI training. By seeking substantial monetary relief, the lawsuit signals that publishers view AI scraping as a direct threat to their revenue streams and intellectual property rights.
Evidence of Systemic Scraping
The complaint moves beyond general allegations to present specific internal data, detailing the scale of the alleged infringement. According to the filing, content from USA Today and its affiliated newspapers comprises more than 160,000 entries within OpenAI’s WebText corpus. This dataset was reportedly used specifically for training the GPT-2 model, suggesting a deep integration of the plaintiffs' journalism into the foundational architecture of the AI.
This volume of data raises critical questions about how the material was acquired. The lawsuit alleges that Microsoft facilitated this collection by providing OpenAI with a copy of the Bing Index. Furthermore, it claims Microsoft operated a crawler known as Project Mango on OpenAI's behalf to gather this content.
- 160,000+ entries from plaintiff publications found in the WebText corpus
- Bing Index reportedly shared by Microsoft to aid data collection
- Project Mango alleged to be a crawler operated for OpenAI
For creators, this evidence suggests that training data may not be gathered through isolated, manual efforts but via systematic, corporate-scale infrastructure. If these allegations hold, it implies that rights holders may have had no practical means to opt out of such automated, large-scale scraping.
The Core Legal Conflict
At the heart of USA Today Co. v. OpenAI is a fundamental clash over intellectual property boundaries. The plaintiffs argue that training large language models on copyrighted journalism constitutes direct infringement, rejecting the defense that such use qualifies as fair use. This dispute centers on whether the transformative nature of AI training justifies the unauthorized reproduction of protected works.
The complaint provides concrete evidence to support this legal stance. It highlights that the plaintiffs' content accounts for more than 160,000 entries in the WebText corpus used for GPT-2 training. Furthermore, the filing includes exhibits of copyright registrations alongside specific examples of GPT-5.6 outputs that allegedly replicate the plaintiffs' journalism.
- Internal data showing extensive corpus inclusion
- Alleged Microsoft assistance via Bing Index and Project Mango
- Specific GPT-5.6 outputs mimicking original articles
For creators, this conflict is critical. If courts rule against fair use, it could fundamentally alter licensing agreements and payment structures, requiring explicit consent for any AI training on existing content.
A Wave of Litigation
The USA Today filing is part of a broader legal surge, with dozens of copyright owners now suing tech giants like OpenAI, Anthropic, and Meta. This wave of litigation signals a fundamental shift in how intellectual property is treated in the AI era, moving from isolated disputes to a systemic challenge against major platforms.
For creators, this trend highlights the fragility of current licensing models. As courts evaluate these cases, the legal landscape is rapidly evolving, potentially redefining what constitutes fair use for training data.
Key aspects of this shifting environment include:
- Multiple high-profile suits targeting different AI developers simultaneously.
- Allegations of unauthorized use of copyrighted material across various platforms.
- A growing body of legal precedent that could influence future industry standards.
These developments suggest that the rights of content owners are becoming a central battleground in the AI industry, with potential long-term impacts on how creators are compensated and protected.
Implications for Creator Contracts
The USA Today lawsuit, seeking over $250 million in damages, signals a critical shift in how platforms must negotiate data rights. As courts evaluate whether unauthorized scraping constitutes infringement, existing licensing agreements may require urgent renegotiation. This legal pressure forces a re-evaluation of how independent professionals and media outlets define their intellectual property boundaries.
For creators, this means contracts must explicitly address AI training usage. Key considerations include:
- Explicit Licensing Terms: Clearly defining whether content is licensed for model training.
- Data Usage Rights: Specifying if data can be used for commercial AI products.
- Compensation Structures: Establishing fair payment models for content inclusion.
According to reports, these clarifications are essential to protect revenue streams. Without precise contractual language, creators risk losing control over their work’s commercial value in the AI era.
Protecting Your Work in the AI Era
Audit Your Digital Footprint
As litigation like USA Today Co. v. OpenAI progresses, creators must proactively audit their digital presence to understand how their work circulates online. This involves identifying where your content is hosted and whether it is easily accessible to automated crawlers. Understanding the specific licensing terms of your platforms is equally critical, as standard agreements may inadvertently permit broad data usage.
To safeguard your intellectual property, consider these actionable steps:
- Review your publishing contracts for clauses regarding AI training and data scraping.
- Check if your content appears in public datasets or AI training corpora.
- Consult legal counsel to interpret "fair use" boundaries in your specific context.
By taking these measures, you can better protect your rights and ensure that your creative output is not exploited without proper compensation or authorization.
FAQ
Why did USA Today sue OpenAI?
USA Today Co. and affiliated newspapers sued OpenAI for allegedly using hundreds of thousands of copyrighted articles to train its large language models without authorization. The complaint cites internal data showing that over 160,000 entries from the plaintiffs' publications were included in the WebText corpus used for GPT-2 training.
How much money is USA Today seeking in the lawsuit?
The plaintiffs are seeking damages exceeding $250 million from OpenAI. They are also requesting a court order to stop the alleged copyright infringement.
What role did Microsoft play according to the lawsuit?
The filing alleges that Microsoft provided OpenAI with a copy of the Bing Index to aid in data collection. Additionally, the complaint states that Microsoft operated a crawler called Project Mango on OpenAI's behalf.
Sources
Draft any contract in minutes — not billable hours
AiDocX generates artist, producer, influencer and crew agreements from a single prompt, then gets them e-signed. Free to start.
Try AiDocX free →Related contract templates
- Free Contract Templates for Creators (hub) →
- Artist Management Agreement Template →
- Music Producer Agreement Template →
- Beat License Agreement Template →
- Music Booking / Performance Agreement →
- Film & Video Crew Agreement Template →
- Influencer–Brand Collaboration Agreement →
- NDA for Creators & Collaborations →