Technology7 min read

Microsoft Exec Warned AI Scraping Is 'Largest Labor Theft'

Unsealed court documents reveal a Microsoft exec privately called AI training on news content 'the largest theft of labor in human history.' Here's what it means.

Microsoft Exec Warned AI Scraping Is 'Largest Labor Theft'

Key takeaways

  1. 1It came from Brent Hecht, Microsoft's Director of Applied Science, in internal documents that the company fought hard to keep sealed.
  2. 2The fourth factor — market harm — has historically carried the most weight in major fair use determinations, including the Supreme Court's 1994 decision in Campbell v.
  3. 3Legal scholars including those affiliated with institutions such as Columbia Law School and Berkeley's Center for Law and Technology have noted publicly that U.
  4. 4What This Means for the Future of AI and Content Creators The internal warnings now on the public record are not simply embarrassing for Microsoft.
Sections · 5

Microsoft Executive's Explosive Internal Warning on AI Scraping

A senior Microsoft scientist tasked with studying the societal impact of AI once wrote that the company's own strategy for training large language models amounted to what he called perhaps the "largest theft of labor in human history." That warning did not come from a critic outside the company, from an academic paper, or from a plaintiffs' attorney. It came from Brent Hecht, Microsoft's Director of Applied Science, in internal documents that the company fought hard to keep sealed.

Those documents are now public. A motion for summary judgment filed by news organizations led by The New York Times was unsealed in federal court, and the excerpts it contains have reframed the ongoing copyright battle over AI training in a way that no press release or earnings call ever could. The question before the court has always been whether companies like Microsoft and OpenAI treated news publishers fairly. The newly unsealed materials suggest that, at minimum, one of Microsoft's own scientists believed the answer was no — and said so in writing, repeatedly.

Hecht's characterization of the scraping program as "an astonishing theft of unprecedented proportions" carries particular weight precisely because it originated inside the organization that built Copilot and bankrolled the development of ChatGPT. These are not the words of a hostile witness. They are the documented concerns of a researcher who understood both the technology and the cultural industries it would disrupt.

The Unsealed Documents: What They Reveal

The Unsealed Documents: What They Reveal — Linkedin login screen with join now option
The Unsealed Documents: What They Reveal — Linkedin login screen with join now option

Court filings are rarely dramatic, but the unsealed motion for summary judgment from the news plaintiffs functions as a curated exhibit of the gap between Microsoft and OpenAI's public messaging and their internal deliberations. According to the news organizations' account of the documents, Hecht returned to the theme of content acquisition and ethical responsibility more than once, framing the mass scraping of news content not as a legally ambiguous edge case but as an ethical crisis hiding in plain sight.

Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026

The sheer scale of what AI companies consumed during the training of models like GPT-4 is difficult to visualize in the abstract. News publishers produce millions of articles per year, each representing hours of reporting, editing, fact-checking, and legal review. Investigative pieces can cost tens of thousands of dollars in staff time alone. When that content is ingested wholesale into training corpora — without compensation, without licensing, without so much as a notification to the producing organization — the economic math becomes stark. Hecht appears to have understood that math and found it troubling enough to put on record.

The documents also exposed internal contradictions in Microsoft and OpenAI's legal posture. While the companies have publicly maintained that training AI on news content constitutes fair use under U.S. copyright law, Hecht's writings told a different story. His internal assessments directly undermined that position, offering the plaintiffs a significant evidentiary foothold as the case moves toward a potential trial.

The Fair Use Debate: A 'Complete Mockery'?

The Fair Use Debate: A 'Complete Mockery'? — Linkedin login screen with join now option
The Fair Use Debate: A 'Complete Mockery'? — Linkedin login screen with join now option

Hecht's most legally significant contribution to the unsealed record may be his framing of the fair use question. According to the news organizations, he wrote that Microsoft and OpenAI's broad approach to scraping news made "a complete mockery of the idea of 'fair use.'" That phrase matters because fair use is the central legal defense both companies have relied upon throughout the litigation.

Under Section 107 of the U.S. Copyright Act, courts evaluate fair use through four factors: the purpose and character of the use, the nature of the copyrighted work, the amount of the work used, and the effect of the use on the market for the original. The fourth factor — market harm — has historically carried the most weight in major fair use determinations, including the Supreme Court's 1994 decision in Campbell v. Acuff-Rose Music. AI companies argue their use is transformative: training a model is not the same as reproducing an article. Critics argue the fourth factor is devastating to that claim, because the output competes directly with the original sources and erodes subscription revenue, referral traffic, and licensing opportunities.

Legal scholars including those affiliated with institutions such as Columbia Law School and Berkeley's Center for Law and Technology have noted publicly that U.S. courts have never squarely applied the four-factor test to large-scale generative AI training, leaving the doctrine in a state of genuine uncertainty. That uncertainty has been the oxygen for the AI industry's legal strategy. Hecht's documents suggest that, internally, at least some Microsoft researchers were not as confident in that uncertainty as the company's lawyers.

The plaintiffs' motion argues that his assessments effectively admit what the litigation seeks to prove: that the companies knew their content acquisition strategy was ethically and legally problematic before they deployed the products that depended on it.

The New York Times-Led Lawsuit and Its Broader Implications

The New York Times filed its copyright lawsuit against Microsoft and OpenAI in late 2023, becoming the highest-profile news organization to take the AI industry to court over training data. The Times was subsequently joined by other outlets in what has grown into a coordinated legal effort by news publishers who argue that decades of journalism were absorbed without consent or compensation into systems that now compete with the very institutions that produced the underlying content.

The unsealing of the summary judgment motion marks a meaningful procedural milestone. Motions for summary judgment ask a court to decide the case — or portions of it — without a full trial, based on the argument that no genuine dispute of material fact exists. The fact that the news plaintiffs are seeking summary judgment in their favor suggests confidence in the documentary record, and the inclusion of Hecht's statements as exhibit-quality evidence indicates they believe those documents speak for themselves.

Beyond the specifics of this case, the litigation has become a proxy war over the terms on which the AI industry will be allowed to develop. A ruling favorable to the Times and its co-plaintiffs could establish that training on copyrighted news content without licensing is not protected by fair use — a precedent that would require AI companies to either negotiate with publishers or fundamentally change how they acquire training data. Either outcome carries enormous financial implications. Licensing the full corpus of English-language journalism that has been scraped would represent a cost structure the current AI business model has not priced in.

What This Means for the Future of AI and Content Creators

The internal warnings now on the public record are not simply embarrassing for Microsoft. They represent a potential inflection point for how regulators, legislators, and courts understand the AI industry's self-awareness about the harm it was causing even as it raced toward commercialization. The knowing-and-proceeding-anyway narrative is one of the most legally and politically damaging frames a company can face.

News organizations are not the only content producers watching this case. Book authors, photographers, screenwriters, software developers, and academic researchers have all raised analogous concerns about their work appearing in training corpora without consent. Several separate copyright suits have been filed by authors and artists in parallel proceedings. The outcome in the Times-led case may set the terms for all of them.

For the AI companies, the disclosure is a reminder that internal communications rarely stay internal forever, particularly in high-stakes commercial litigation where both sides have strong incentives to surface damaging documents. Microsoft and OpenAI built their AI products at extraordinary speed. That speed appears to have outpaced any serious internal reckoning with what Hecht plainly saw: that building a new industry on top of another industry's labor, without asking or paying, is not a legal gray area. It is, in his words, theft of unprecedented proportions.

Whether the courts ultimately agree will determine not just the fate of these lawsuits, but the foundational rules governing how artificial intelligence is allowed to learn.


Source: AI - Ars Technica

Published

29 September 2026

Author

Editorial

Comments

No comments yet. Be the first.

Leave a comment