Advertisement

The AI Copyright War Just Got Real – OpenAI Ordered to Reveal Deleted Evidence

December 2nd, 2025Jump to Comment Section3
The AI Copyright War Just Got Real – OpenAI Ordered to Reveal Deleted Evidence

A federal judge has ordered OpenAI to reveal internal communications about why it deleted two massive datasets of pirated books, a ruling that could expose whether the company knowingly built its products on stolen creative work. For filmmakers whose content may have been scraped without permission, this case establishes critical legal precedent about AI training practices.

The November 2025 ruling by U.S. Magistrate Judge Ona T. Wang requires OpenAI to hand over Slack messages and attorney communications explaining why it destroyed datasets called “Books1” and “Books2” containing over 100,000 books downloaded from the pirate library LibGen. The decision arrives just months after competitor Anthropic settled similar claims for a record 1.5 billion dollars, as we reported in our coverage of Sora 2’s copyright reckoning.

While this case involves text, the legal principle matters for all creators: courts are finding that obtaining copyrighted content from pirate sources is “inherently, irredeemably infringing” regardless of downstream use. The same framework applies to films, cinematography, and visual effects harvested from torrent sites or pirate platforms.

Sam Altman speaking at TED. Image credit: Steve Jurvetson, license CC BY 2.0

What the court is forcing OpenAI to reveal

Judge Wang’s ruling centers on datasets an OpenAI employee downloaded from Library Genesis in 2018, later deleted in 2022, one year before any lawsuits were filed. These are the only training datasets OpenAI has ever deleted.

The company must now produce communications from Slack channels named “project-clear” and “excise-libgen” where employees discussed the deletion. OpenAI initially claimed the datasets were removed “due to non-use,” then attempted to withdraw that statement and assert attorney-client privilege. The court rejected this approach, finding that OpenAI waived privilege by making inconsistent claims.

If communications reveal the company knew it was using pirated material, statutory damages jump from 30,000 dollars to 150,000 dollars per work. With hundreds of thousands of works potentially involved, the financial exposure reaches billions.

The “project-clear” Slack channel. Illustration credit: CineD

Why filmmakers should pay attention

The Anthropic settlement validates the legal theory now targeting OpenAI. In September 2025, Anthropic paid 1.5 billion dollars after a judge ruled that downloading 7 million pirated books made all downstream use infringing. Lead counsel called it “the largest publicly reported copyright recovery in history.”

Hollywood studios have launched their own actions. As we covered when Disney and NBCUniversal sued Midjourney, major entertainment companies are calling AI image generators “a bottomless pit of plagiarism.” In November 2025, Disney and Warner Bros. Discovery jointly sued Chinese AI company MiniMax over its Hailuo AI video tool.

For cinematographers, the question isn’t just whether AI can reproduce your specific shots. It’s whether the AI company legally acquired the material it learned from. Courts are increasingly answering no.

The legal landscape is shifting toward creators

Several developments favor content creators. A German court issued the first European ruling against OpenAI in November 2025 for reproducing song lyrics. The U.S. Copyright Office found that arguments AI training is “inherently transformative” are mistaken. And the Andersen v. Stability AI case, heading to trial in September 2026, could establish that AI models containing unlicensed training data are themselves infringing works.

This matters for practical tool choices. Some companies are responding by licensing content properly. As we noted in our coverage of Moonvalley’s Marey model, “commercially safe” AI trained on licensed material is emerging as an alternative to legally questionable tools. Drew Geraci’s new ethical AI video course on MZed addresses these concerns directly.

OpenAI has appealed the discovery ruling and moved to pause enforcement. The consolidated litigation has a discovery deadline of February 2026, with no final ruling on fair use expected until summer 2026.

The OpenAI case will help determine whether filmmakers have real recourse when their work is used to train AI without permission. As I discussed in my BILD Expo presentation on AI workflows, generative AI has been trained on massive datasets of copyrighted work, and none of us gave permission.

Have you discovered your cinematography in AI training datasets or outputs? Don’t hesitate to let us know in the comments below!

3 Comments

Filter:
all
Sort by:
latest
Filter:
all
Sort by:
latest