TL;DR — Key Takeaways

– Unredacted court filings reveal internal concerns at Microsoft and OpenAI about the impact of AI products on publishers and copyrighted content.

– Microsoft’s Brent Hecht warned internally that AI scraping could be viewed as an unprecedented theft of creators’ work and could undermine the publishers AI systems depend on.

– Microsoft data cited by the plaintiffs showed Copilot generated significantly fewer referral clicks to New York Times content than traditional Bing search.

Unredacted court filings in The New York Times’ landmark copyright lawsuit against OpenAI and Microsoft Corp. have exposed stark contradictions between the tech giants’ public legal defenses and their internal assessments: Some top executives privately describe artificial intelligence (AI) web scraping as the “largest theft of labor in human history.”

The filings, surfaced Thursday in support of the plaintiffs’ motion for summary judgment, detail internal communications and sworn depositions from key leadership across both companies.

The New York Times, joined by the New York Daily News and the Center for Investigative Reporting, alleges OpenAI and Microsoft illegally used millions of copyrighted news articles to train their generative AI models.

In a January 2023 internal memo, Brent Hecht, Microsoft’s director of applied science, wrote that millions of people would view AI models “hoovering up” their work as “an astonishing theft of unprecedented proportions.” Hecht noted that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”

Hecht later warned of a self-inflicted “doom loop” in a January 2024 presentation, writing that it is “highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created.”

Internal data cited in the filings revealed that Microsoft Copilot’s answer engine produced up to 93% fewer referral clicks to The New York Times domain compared to traditional Bing search, systematically starving the newsroom of vital traffic and weakening the very content supply chain AI relies on.

Concerns extended to OpenAI’s executive ranks. Nick Turley, head of ChatGPT, internally characterized AI products as an “existential threat” to publishers, admitting the tools are “largely substitutive” and will become increasingly so as the technology improves. Another OpenAI engineer testified that “no matter how prominently we show the links, users won’t click.”

The documents also allege intentional bypasses of digital subscriptions. In one exchange, OpenAI researcher Nick Ryder informed company president Greg Brockman of a “hack to get around [the] nytimes paywall,” to which Brockman reportedly replied, “ah nice.” Additionally, a joint initiative dubbed Project Mango allegedly assembled a dataset containing at least 160,903 unique works from the plaintiff publishers, with copyright notices deliberately stripped prior to training.

Under oath, Microsoft CEO Satya Nadella appeared to distance himself from such practices.

In his deposition, Nadella agreed that chatbot interactions substitute for visiting original publisher sites, testifying that “anything that is paywalled should be licensed.” Nadella said he would have required OpenAI to retrain its models had he known paywalled material was used without permission.

While OpenAI and Microsoft publicly maintain that scraping public web content constitutes non-infringing fair use, arguing the process is transformative, legal experts suggest these admissions could severely undercut that defense.

A key factor in evaluating fair use is the effect on the original work’s market value. By acknowledging internally that their products create a direct substitute for original journalism and erode publisher revenues, both tech companies may face an uphill battle in court.

Frequently Asked Questions

What is The New York Times’ lawsuit against OpenAI and Microsoft about?
The Times alleges the companies used copyrighted journalism without authorization to train and operate generative AI systems.
What did the newly unredacted filings reveal?
They disclosed internal Microsoft and OpenAI communications discussing AI scraping, publisher traffic losses, paywalled content and the potential threat AI products pose to news organizations.
What did Microsoft’s Brent Hecht say?
Hecht wrote that many creators could view AI companies taking their work for training as an extraordinary form of theft and warned that AI could damage the economic foundation of its content suppliers.