The phrase "doom loop" rarely appears in corporate documentation without someone, somewhere, understanding the gravity of what they are describing. When recently unsealed court filings in the New York Times' lawsuit against OpenAI and Microsoft surfaced language that stark, it did not come from outside critics or alarmed regulators. It came from the companies themselves.
Discovery documents unsealed as part of ongoing copyright litigation show that OpenAI and Microsoft were aware, in their own written communications, that their approach to building AI systems carried consequences for the broader web — consequences they apparently named and documented before the public conversation about those consequences had fully formed.
What the Unsealed Court Documents Reveal
Internal documentation produced during discovery in the New York Times case contains two characterizations that have drawn immediate attention from legal scholars and industry observers alike. First, the companies' own materials warned that AI training practices could initiate a "doom loop" for the web — a self-reinforcing cycle of degradation in which AI systems consume content, reduce the incentive to create new content, and thereby diminish the quality of training data available for future models. Second, the same documentation described the mass scraping of web content to train AI models as the "largest theft of labor in human history."
These are not reconstructions by opposing counsel. They are phrases that originated within the organizations defending against the Times' claims.
The significance of this distinction cannot be overstated. In copyright litigation, fair use arguments typically depend on demonstrating that a defendant acted in good faith, that the use was transformative, and that it did not harm the market for original works. Internal admissions that a practice constitutes theft — and that it would damage the ecosystem from which the data was drawn — directly complicate each of those prongs. Legal scholars who study digital copyright have long noted that a defendant's contemporaneous awareness of potential harm weighs heavily against fair use claims, particularly the fourth factor, which examines market harm to the original works.
The New York Times Case Against OpenAI and Microsoft
The New York Times filed its lawsuit against OpenAI and Microsoft in late 2023, alleging that the companies used millions of the publication's articles without authorization to train large language models including ChatGPT. The case is among the most closely watched in a wave of AI copyright litigation that also involves lawsuits from book authors, record labels, and visual artists.
Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026What makes the Times action particularly consequential is its scale and the specificity of its claims. Unlike some plaintiffs who have struggled to demonstrate direct, quantifiable harm, the Times has argued that AI systems can reproduce substantial portions of its journalism nearly verbatim — undercutting the market for licensed content and potentially redirecting readers who might otherwise visit NYTimes.com to interact with AI summaries instead.
The unsealing of discovery materials is a procedural development that often shifts litigation dynamics. When documents previously shielded under protective orders become part of the public record, they can reshape settlement calculations, inform parallel regulatory proceedings, and — as appears to be happening here — generate significant press coverage that creates reputational pressure independent of the legal outcome.
Microsoft's exposure in this case stems from its multi-billion-dollar investment in OpenAI and its deep integration of OpenAI's models into products including Bing, Copilot, and Azure services. The companies have mounted a joint defense anchored in the argument that training AI on publicly available web data constitutes fair use under U.S. copyright law.
How AI Scraping Threatens the Open Web
The "doom loop" framing in the unsealed documents maps onto a pattern that web publishers have been reporting with increasing alarm. AI-powered search features, including Google's AI Overviews, have corresponded with measurable declines in referral traffic to news sites and other content publishers. Analyses from firms including Similarweb and Cloudflare have tracked shifts in how users interact with search results when AI summaries appear — users who receive a satisfactory answer within a search interface have less reason to click through to the source.
The mechanism is straightforward: publishers produce content that gets scraped to train models, those models then answer questions that previously would have driven users to publisher sites, publishers see declining traffic and revenue, and the economic foundation for producing the content that feeds future model training erodes. The loop closes on itself.
This is not a hypothetical future scenario. Independent web publishers have reported traffic declines ranging from 20 to over 60 percent in the period following the broad rollout of AI search features, depending on their vertical and content type. Reference-heavy topics — health information, legal guides, product reviews — have been hit disproportionately hard, as these are precisely the query types where AI summaries can substitute for a full article visit.
The characterization of scraping as the "largest theft of labor in human history" in OpenAI and Microsoft's own documents points to a scale argument. Web content represents decades of work by journalists, researchers, educators, developers, and countless other contributors who published under the assumption that readers, not training pipelines, were their audience. The aggregate value of that labor, captured without compensation or consent, is genuinely difficult to quantify — but the companies themselves appear to have attempted some version of that calculation.
What This Means for the Future of AI and Content
The unsealed materials put pressure on a negotiated settlement path that has been quietly developing in parallel to litigation. Several major publishers, including Associated Press and some regional newspaper chains, have entered licensing agreements with AI companies. The existence of a licensing market is legally relevant: it suggests there is an established commercial pathway for obtaining content rights, which in turn supports the argument that bypassing that pathway causes market harm.
Copyright scholars have observed that the AI industry's fair use defense faces a structural difficulty that internal documents of this nature make harder to sustain. The transformative use argument — that training an AI model is sufficiently different from reproducing the original work — is more persuasive when the defendant can demonstrate no awareness that the underlying activity was harmful. Documents showing contemporaneous concern about "doom loops" and "theft" suggest awareness of harm, not good-faith uncertainty.
Courts have not yet definitively ruled on AI training and fair use in any major case. The New York Times litigation, along with parallel cases involving book authors represented by the Authors Guild, will likely produce precedents that govern the industry for years.
Industry and Regulatory Implications
The political and regulatory context has shifted considerably since these cases were filed. In the United States, the Copyright Office has been actively soliciting comment on AI training data, and Congressional interest in amending copyright frameworks to address AI has grown on a bipartisan basis. In the European Union, the AI Act's provisions on training data transparency are already in effect for the largest model developers.
For OpenAI and Microsoft specifically, the unsealed documents arrive at a moment when both companies are navigating significant public and regulatory scrutiny. OpenAI's transition away from its nonprofit structure and Microsoft's continued defense of its AI investments have attracted attention from attorneys general, competition regulators, and institutional investors concerned about governance and liability exposure.
The practical question for the industry is whether voluntary licensing frameworks can develop quickly enough to forestall mandatory regulatory intervention. Several AI developers have signaled openness to licensing agreements, but the terms — particularly the question of retroactive compensation for data already used in training — remain deeply contested.
What the New York Times litigation has now placed on the public record is something that had long been suspected but not confirmed: at least some AI developers understood the structural risks of their data acquisition practices and documented those risks internally. Whether that understanding translates into legal liability, or spurs a broader industry reckoning with how AI systems source the knowledge they transmit, is a question that courts, regulators, and the market will now have to answer together.
Source: The Verge



