Unredacted Filings Detail Internal Microsoft Warning of AI 'Theft of Labor'

Newly unredacted legal filings in the ongoing copyright infringement litigation against Microsoft and OpenAI have exposed damning internal admissions from executives at both technology giants. In internal communications cited in court documents, senior technical leaders privately described web scraping practices as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history," directly contradicting public assertions that training Large Language Models (LLMs) on copyrighted news reporting constitutes fair use under federal law.
The documents, submitted as part of briefs in the lawsuit brought by The New York Times, the Daily News, and the Center for Investigative Reporting, outline how generative artificial intelligence platforms actively erode the economic foundations of media companies. Internal analytics compiled by Microsoft revealed that its Copilot search tool caused click-through rates to The New York Times' primary web domain to plunge by as much as 93% compared to standard Bing search queries—a systemic drop that internal researchers categorized as a self-destructive cycle for the open web.
The filings also detail high-level executive awareness of paywall circumvention and show deep integration between Microsoft and OpenAI's content acquisition pipelines. As legal scrutiny intensifies, these unredacted excerpts present an acute challenge to the defense strategy relied upon by AI developers, providing plaintiffs with internal evidence that AI models act as direct market substitutes rather than transformative tools.
Key Developments & Policy Breakdown
- Internal Warning on Systemic Scraping: In a January 2023 internal memo, Microsoft's Director of Applied Science, Brent Hecht, characterized automated AI data extraction as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."
- Devastating Impact on Referral Traffic: Internal Microsoft analytics from January 2024 revealed that the company's Copilot "answer engine" reduced click-through traffic to The New York Times by up to 93% relative to standard search engines, creating what Hecht termed a "doom loop" that degrades both model performance and publisher revenues simultaneously.
- Direct Subversion of Paywalls: Filings detail an internal exchange where OpenAI researcher Nick Ryder informed President Greg Brockman of a technical "hack to get around nytimes paywall," to which Brockman allegedly responded, "ah nice," undermining claims of good-faith data acquisition.
- Massive Dataset Aggregation: Court filings show OpenAI’s mid-training datasets contained over 91,692 copies of works from the plaintiff news outlets, while a Common Crawl dataset amassed more than 2 million documents from nytimes.com alone.
- Joint Enterprise Initiatives: Microsoft and OpenAI operated joint data sharing pipelines, designated internally as "Project Taxi" and "Project Mango," the latter assembling a training repository holding at least 160,903 unique copyrighted publisher works.
- Deposition Testimony Under Oath: Microsoft CEO Satya Nadella testified that paywalled material ought to be licensed by anyone training AI models, stating under oath that had he known OpenAI scraped paywalled content, he would have exercised Microsoft's rights to require OpenAI to retrain its models.
In-Depth Analysis & Real-World Impact
The central legal defense for generative AI developers rests on the doctrine of fair use, specifically arguing that model training transforms raw text into entirely new algorithmic capabilities without harming the commercial market for the original works. However, the unredacted filings directly target the fourth factor of the statutory fair use test: the effect of the use upon the potential market for or value of the copyrighted work. By documenting that AI chatbots act as "largely substitutive" answer engines—delivering complete reporting directly inside the chat interface—publishers argue that tech companies are destroying the referral-based business model of the open web.
The internal concept of the "doom loop" highlights a broader economic vulnerability for the AI sector itself. As answer engines capture user attention and starve newsrooms of direct web traffic and digital advertising revenues, publishers face severe budget cuts or insolvency. This depletion of original source material inevitably degrades the quality of future AI training datasets, creating a dynamic where LLMs destroy the very ecosystem required to sustain their factual accuracy and real-time utility.
Furthermore, the revelation of joint projects like Project Mango and Project Taxi complicates Microsoft’s attempts to distance itself from OpenAI's operational choices. By showing that Microsoft actively evaluated, processed, and exchanged scraped content to enhance its commercial Copilot product, the filings strengthen the plaintiffs' claims of vicarious and contributory copyright infringement against the enterprise tech titan.
Background, Preceding Events & Historical Context
The legal clash traces back to late 2023 when The New York Times filed a landmark copyright suit against OpenAI and Microsoft, alleging that millions of its articles were ingested without authorization to build commercial AI platforms. Over the following months, the litigation expanded as other media organizations, including the Center for Investigative Reporting and Alden Global Capital's MediaNews Group, launched parallel legal actions. The tech sector had historically operated under implicit web scraping conventions established during the search engine era, where search crawlers index web pages in exchange for driving direct referral traffic back to creators.
The rise of generative AI fundamentally broke this social contract. Unlike traditional search engines that direct users outward to source material via hyperlinks, LLMs digest raw text to synthesize direct answers within proprietary interfaces. The unredacted filings reveal that behind closed doors, leaders at both Microsoft and OpenAI recognized this structural shift early on. Communications from Nick Turley, head of ChatGPT, acknowledged that chatbots were increasingly substitutive, while OpenAI President Greg Brockman noted that the models had become "excellent at news," validating publishers' arguments that AI outputs directly compete with primary journalism.
“"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain." — Internal Microsoft Presentation, January 2024”
Strategic Outlook & What to Watch Next
As the case moves through pre-trial discovery and motions for summary judgment, the evidentiary record revealed in these unredacted briefs will likely force a major re-evaluation of legal strategies across the artificial intelligence sector. If federal courts determine that internal acknowledgments of market harm and deliberate paywall circumvention invalidate fair use protections, Microsoft and OpenAI could face statutory damages reaching billions of dollars, alongside potential court orders mandating the destruction or costly retraining of foundational models built on scraped data.
Moving forward, industry analysts should monitor whether these revelations accelerate enterprise-wide licensing agreements between AI developers and legacy media conglomerates. Facing potential judicial precedents that restrict automated web scraping, foundational model builders may be forced to abandon non-consensual data collection in favor of standardized commercial licensing frameworks—a shift that would dramatically increase the capital required to build next-generation AI software.
Quik News synthesizes verified facts across international press reporting. Original reporting belongs to the attributed outlets above.




