NEW PODCAST: IBC 2026 & Best-of-Show Winners | iPhone 18 Pro Gets a Variable Aperture – Focus Check ep134 →🎙️ WATCH/LISTEN Now
Watch/Listen Now IBC 2026 & iPhone 18 Pro🎙️NEW PODCAST
Education for Filmmakers
Language
The CineD Channels
Info
New to CineD?
You are logged in as
We will send you notifications in your browser, every time a new article is published in this category.
You can change which notifications you are subscribed to in your notification settings.
A federal judge has ordered OpenAI to reveal internal communications about why it deleted two massive datasets of pirated books, a ruling that could expose whether the company knowingly built its products on stolen creative work. For filmmakers whose content may have been scraped without permission, this case establishes critical legal precedent about AI training practices.
The November 2025 ruling by U.S. Magistrate Judge Ona T. Wang requires OpenAI to hand over Slack messages and attorney communications explaining why it destroyed datasets called “Books1” and “Books2” containing over 100,000 books downloaded from the pirate library LibGen. The decision arrives just months after competitor Anthropic settled similar claims for a record 1.5 billion dollars, as we reported in our coverage of Sora 2’s copyright reckoning.
While this case involves text, the legal principle matters for all creators: courts are finding that obtaining copyrighted content from pirate sources is “inherently, irredeemably infringing” regardless of downstream use. The same framework applies to films, cinematography, and visual effects harvested from torrent sites or pirate platforms.
Judge Wang’s ruling centers on datasets an OpenAI employee downloaded from Library Genesis in 2018, later deleted in 2022, one year before any lawsuits were filed. These are the only training datasets OpenAI has ever deleted.
The company must now produce communications from Slack channels named “project-clear” and “excise-libgen” where employees discussed the deletion. OpenAI initially claimed the datasets were removed “due to non-use,” then attempted to withdraw that statement and assert attorney-client privilege. The court rejected this approach, finding that OpenAI waived privilege by making inconsistent claims.
If communications reveal the company knew it was using pirated material, statutory damages jump from 30,000 dollars to 150,000 dollars per work. With hundreds of thousands of works potentially involved, the financial exposure reaches billions.
The Anthropic settlement validates the legal theory now targeting OpenAI. In September 2025, Anthropic paid 1.5 billion dollars after a judge ruled that downloading 7 million pirated books made all downstream use infringing. Lead counsel called it “the largest publicly reported copyright recovery in history.”
Hollywood studios have launched their own actions. As we covered when Disney and NBCUniversal sued Midjourney, major entertainment companies are calling AI image generators “a bottomless pit of plagiarism.” In November 2025, Disney and Warner Bros. Discovery jointly sued Chinese AI company MiniMax over its Hailuo AI video tool.
For cinematographers, the question isn’t just whether AI can reproduce your specific shots. It’s whether the AI company legally acquired the material it learned from. Courts are increasingly answering no.
Several developments favor content creators. A German court issued the first European ruling against OpenAI in November 2025 for reproducing song lyrics. The U.S. Copyright Office found that arguments AI training is “inherently transformative” are mistaken. And the Andersen v. Stability AI case, heading to trial in September 2026, could establish that AI models containing unlicensed training data are themselves infringing works.
This matters for practical tool choices. Some companies are responding by licensing content properly. As we noted in our coverage of Moonvalley’s Marey model, “commercially safe” AI trained on licensed material is emerging as an alternative to legally questionable tools. Drew Geraci’s new ethical AI video course on MZed addresses these concerns directly.
OpenAI has appealed the discovery ruling and moved to pause enforcement. The consolidated litigation has a discovery deadline of February 2026, with no final ruling on fair use expected until summer 2026.
The OpenAI case will help determine whether filmmakers have real recourse when their work is used to train AI without permission. As I discussed in my BILD Expo presentation on AI workflows, generative AI has been trained on massive datasets of copyrighted work, and none of us gave permission.
Have you discovered your cinematography in AI training datasets or outputs? Don’t hesitate to let us know in the comments below!
Δ
Stay current with regular CineD updates about news, reviews, how-to’s and more.
You can unsubscribe at any time via an unsubscribe link included in every newsletter. For further details, see our Privacy Policy
Want regular CineD updates about news, reviews, how-to’s and more?Sign up to our newsletter and we will give you just that.
You can unsubscribe at any time via an unsubscribe link included in every newsletter. The data provided and the newsletter opening statistics will be stored on a personal data basis until you unsubscribe. For further details, see our Privacy Policy
Nino Leitner, AAC is Co-CEO of CineD and MZed. He co-owns CineD (alongside Johnnie Behiri), through his company Nino Film GmbH. Nino is a cinematographer and producer, well-traveled around the world for his productions and filmmaking workshops. He specializes in shooting documentaries and commercials, and at times a narrative piece. Nino is a studied Master of Arts. He lives with his wife and two sons in Vienna, Austria.