Court Docs: AI Execs Knew Chatbots Threatened Journalism
A newly unredacted court filing quotes Microsoft and OpenAI insiders calling AI training data use an "astonishing theft" and an "existential threat" to publishers, undercutting fair use defenses.

Updated
Why it matters
- Microsoft's Brent Hecht called AI training on copyrighted text "an astonishing theft of unprecedented proportions" in comments cited by publishers.
- An OpenAI executive wrote publishers faced an "existential threat" from chatbots; Greg Brockman replied "ah nice" to a described NYT paywall workaround.
- Microsoft CEO Satya Nadella testified paywalled content "should be licensed" and that he would have required OpenAI to retrain models trained on it.
Microsoft's director of applied science, Brent Hecht, privately called AI training on publishers' text "an astonishing theft of unprecedented proportions" and possibly the "largest theft of labor in human history" — according to quotes cited by news publishers in a court filing unredacted this week.
The brief was filed earlier this month in consolidated copyright litigation against OpenAI and Microsoft in a New York federal court. The proceeding combines related infringement suits, including cases brought by The New York Times and by Ziff Davis, the publisher of CNET.
The publishers' filing cites comments from inside both companies that appear to undercut the defense that training large language models on copyrighted material qualifies as fair use. Under that doctrine, courts weigh how copyrighted works are used, the nature of the work, how much of it was used, and the effect on the market for that work. Exhibits showing the full context of Hecht's quotes remain under seal.
Microsoft CEO Satya Nadella testified that chatbot conversations give users information "right there on the website on the AI platform versus needing to go to the underlying source" — meaning the publisher that originally reported it. An OpenAI executive, in a comment cited by the publishers, wrote that publishers faced an "existential threat" from products like the company's chatbot.
The filing also describes how AI developers worked around paywalls, including The New York Times'. An OpenAI employee told company president Greg Brockman about "a hack to get around nytimes paywall," to which Brockman replied, "ah nice."
Yet the same document records testimony cutting the other way. Nadella said under oath that anything behind a paywall "should be licensed by anyone who wants to use it" for AI development, and that had he known OpenAI trained on paywalled content, he would have required the company to retrain its models.
A Microsoft spokesperson told CNET that Hecht's comments "reflect one employee's perspectives" and that the company's position rests in its court filings, which "explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism." Nadella's remarks, the spokesperson said, addressed changes in how people consume information and are "perfectly consistent" with the company's legal position. "Those observations should not be confused with conclusions about copyright questions before the Court, which Microsoft addresses in its filings."
In its own brief, Microsoft argued that training LLMs on published content is significantly transformative. "Copyright law does not permit rightsholders to block transformative technologies like LLMs; it encourages such uses on the expectation that rightsholders will adapt and the public will be better off for it," the company wrote.
Representatives for OpenAI and Ziff Davis did not immediately respond to requests for comment.
The stakes extend beyond this case. Publishers argue AI companies must pay to license the content they use, while the tech companies counter that licensing requirements are unnecessary and would burden development of more capable AI. The Trump administration weighed in earlier this month, arguing in a statement that a licensing requirement would impede American AI development in the race with China. Related litigation from book authors could compound the consequences for both industries.
The publishers also contend that OpenAI and Microsoft employees understood their chatbots' ability to find, copy, summarize and sometimes regurgitate publisher content was diverting readers away from news sites. "Defendants' own experts acknowledge that grounded LLMs exploit their sources rather than promote them like search uses that have been deemed fair use," the filing states.
How the court weighs these internal admissions against the fair use defense could shape whether AI developers must pay for the training data that underpins their products — and whether news organizations see compensation for the work those models were built on.
Original: courtlistener.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles