OpenAI Accuses Plaintiffs’ Lawyers of Concealing Paid Research in AI Copyright Suit
OpenAI alleges plaintiffs’ counsel secretly funded and concealed research in pending AI copyright lawsuit.
Why it matters: This dispute highlights emerging challenges in managing evidence and litigation strategy in AI copyright cases, potentially affecting court rulings on fair use and discovery processes.
- OpenAI and Microsoft claim Susman Godfrey paid for research outside formal expert disclosures in the AI copyright case.
- The research, involving Jane Ginsburg, argues generative AI harms book markets but faces criticism for flawed data.
- Plaintiffs accuse OpenAI of deleting billions of ChatGPT outputs and withholding internal data, violating preservation orders.
- Microsoft’s documents reveal concerns over AI scraping’s impact on content creators and OpenAI’s bypass of news paywalls.
OpenAI and Microsoft have accused the plaintiffs’ law firm, Susman Godfrey, of secretly funding research submitted as evidence in their ongoing AI copyright lawsuit and concealing those payments from the court and defense counsel.
The contested research, co-authored by scholar Jane Ginsburg and led by Tuhin Chakrabarty, claims generative AI technology harms book markets. However, independent analyst Thad McIlroy criticized this study for relying heavily on Kindle Unlimited data, which may not accurately represent the broader publishing ecosystem.
OpenAI and Microsoft argue that Susman Godfrey introduced this commissioned work outside the court-ordered expert disclosure process, raising ethical and procedural concerns about evidence transparency and fairness in discovery.
In response, plaintiffs allege OpenAI deleted billions of ChatGPT-generated outputs after the start of the lawsuit and withheld crucial internal evaluations and training data search tools, potentially violating court preservation orders. The preservation obligation requires parties to retain relevant evidence, including electronic communications and data, from the time litigation begins.
Reports also indicate OpenAI maintained a separate database of 78 million anonymized ChatGPT conversations used to evaluate potential copyright infringement risks. They developed a detection system internally called the "Bloom" filter, part of "Project Giraffe," designed to identify ChatGPT outputs closely matching training data shortly after litigation commenced.
Internal Microsoft documents, including statements from executive Brent Hecht, described AI web scraping as "the largest theft of labor in human history," reflecting concerns that AI training practices undercut content creators' revenue streams. OpenAI executives have also acknowledged their models’ ability to bypass news paywalls, a practice contested by publishers as challenging fair use defenses.
Lead plaintiffs’ counsel Ian B. Crosby said, "If OpenAI truly believed using our clients’ journalism was lawful, it would not have concealed evidence about those actions." OpenAI spokesperson Drew Pusateri dismissed the claims as "blatantly false allegations," describing them as an attempt to distract from weaknesses in the plaintiffs’ case.
This dispute underscores the evolving legal landscape around AI-generated content, spotlighting critical questions about evidence disclosure, the scope of fair use, and the responsibilities of litigants in digital discovery. As courts contend with novel AI litigation issues, outcomes here could shape future standards for transparency and technology accountability.
By the numbers:
- 78 million — anonymized ChatGPT conversations stored by OpenAI for copyright evaluation
- Billions — ChatGPT outputs plaintiffs say were deleted post-lawsuit start
- Single research study — challenged for reliance on Kindle Unlimited data
Yes, but: Plaintiffs' claims about evidence deletion and concealment are contested by OpenAI, which calls them false and motivated by litigation tactics.
What's next: The case continues to progress in court, with discovery disputes and motions over evidence handling expected to influence procedural rules in AI copyright litigation.