Microsoft's Own Scientist Called AI Scraping 'Largest Theft of Labor in History'
Unsealed court filings show Microsoft's Brent Hecht warned internally that AI news scraping was the "largest theft of labor in human history," undercutting the fair use defense.

Updated
Why it matters
- Documents unsealed Thursday in the news organizations' copyright lawsuit against Microsoft and OpenAI quote Microsoft Director of Applied Science Brent Hecht calling AI news scraping possibly the "largest theft of labor in human history."
- In another internal document, Hecht suggested the plan to widely scrape news made "a complete mockery of the idea of 'fair use,'" contradicting the companies' courtroom defense.
- The documents come from a motion for summary judgment filed by news plaintiffs led by The New York Times, exposing internal communications Microsoft and OpenAI had fought for years to keep confidential.
A Microsoft director of applied science repeatedly warned internally that scraping news content to train AI systems was perhaps the "largest theft of labor in human history," according to documents unsealed Thursday in a copyright lawsuit brought by news organizations against Microsoft and OpenAI.
The newly public filings come from a motion for summary judgment filed by news plaintiffs led by The New York Times. The documents were unsealed on Thursday, and they expose internal communications that the news organizations alleged show exactly how Microsoft and OpenAI viewed the threat to journalism before releasing AI products such as ChatGPT and Copilot.
The central figure in the disclosures is Brent Hecht, Microsoft's Director of Applied Science. In internal documents cited by the plaintiffs, Hecht described the scraping of news for AI training as "an astonishing theft of unprecedented proportions." In another document, Hecht directly undermined the legal argument that Microsoft and OpenAI have mounted publicly in court — that training AI models on copyrighted news content constitutes fair use. According to the news organizations, Hecht suggested that the plan to widely scrape news content made "a complete mockery of the idea of 'fair use.'"
The quotes carry weight because of who uttered them. Hecht is not an outside critic, a competing publisher, or a regulator. He is a senior technical figure inside Microsoft, the company standing as a co-defendant in one of the most consequential copyright cases now working through the US courts. When a company's own applied science director characterizes its data acquisition practices in the language of theft, plaintiffs gain a piece of evidence that is difficult to dismiss as self-interested advocacy from the news industry.
The context matters. For years, according to the source reporting, Microsoft and OpenAI have fought to keep certain information out of the public eye in their legal battle with news organizations. Those organizations accuse the two AI firms of effectively teaming up to violate copyright laws by harvesting large volumes of news content to train their models. The Thursday unsealing marks a turning point in that effort to maintain confidentiality: details that plaintiffs argued should never have been marked confidential are now entering the public record.
At the heart of the dispute is a question with far-reaching commercial and legal stakes: whether AI developers can freely use copyrighted news articles as training data under the doctrine of fair use, or whether they must license that content. The outcome will shape how AI companies build the next generation of models and whether news publishers receive compensation when their reporting becomes raw material for products like ChatGPT and Copilot — products that answer user questions directly and, publishers argue, divert the traffic and revenue that traditionally sustained journalism.
The unsealed documents sharpen that dispute in a specific way. Microsoft and OpenAI have argued in court that training on news content is fair use. Hecht's internal assessment, as characterized by the plaintiffs, suggests that at least one senior Microsoft scientist believed the scale and method of the scraping operation made that position untenable on its own terms. A plan to scrape news broadly, he reportedly suggested, turned the fair use defense into something absurd.
For the news plaintiffs, led by The New York Times, the disclosure serves a strategic function at this stage of the litigation. A motion for summary judgment asks the court to rule on the case without a full trial, and evidence that a defendant's own scientists described the conduct as theft strengthens the plaintiffs' argument that there is no genuine dispute of material fact. Internal admissions, when they exist, are among the most potent assets in copyright litigation, because they undercut a defendant's characterization of its own practices.
The filing also pulls back the curtain on what Microsoft and OpenAI knew, and when. The news organizations alleged that the internal documents show how the two companies assessed the danger to news organizations before shipping AI products to the public. That framing matters for the broader public debate: if AI companies privately recognized that their training practices threatened the economics of news publishing while publicly defending those practices as lawful and benign, the gap between internal assessment and external messaging becomes part of the legal and political record.
Microsoft and OpenAI have not, in the source material, publicly responded to the specific newly unsealed quotes from Hecht. Their formal position in the litigation remains that training AI models on news content falls within fair use — the very position Hecht's internal comments appear to contradict.
The unsealing arrives at a moment when courts, lawmakers, and regulators are still establishing the rules for AI training data. Copyright lawsuits against AI developers have multiplied, and each disclosure of internal deliberations adds to the evidentiary record that will inform how those cases are decided — and, by extension, how the business model of large-scale AI development is structured going forward. If courts side with publishers, licensing costs could reshape the economics of foundation models; if the AI companies prevail, publishers will have few remaining levers to control how their content is used.
The Hecht documents ensure that this case will be watched not only for its legal outcome but for what it reveals about the internal culture of the companies building frontier AI. A Microsoft scientist calling the industry's core data practice the "largest theft of labor in human history" is a line that will follow both defendants through the remainder of the litigation — and into any settlement negotiations, licensing talks, or legislative debates that follow.
As the summary judgment motion moves forward, the key question becomes whether the court treats these internal warnings as evidence that the fair use defense cannot survive scrutiny, or as stray internal commentary that does not settle the legal question. Either way, the confidentiality wall Microsoft and OpenAI built around these documents has now cracked, and the plaintiffs have shown there may be more behind it.
Original: cdn.arstechnica.net
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
144 articles
Related articles
- Court Docs: AI Execs Knew Chatbots Threatened Journalism
- OpenAI Bans Accounts Reviving Russia's 'Stop News' Influence Operation
- Guardian columnist: shame over AI use misses the real target
- AI Models Keep Cheating on Tests, and Researchers Are Quitting
- OpenAI Brings ChatGPT Edu to Journalism Schools at CUNY and Northwestern