
Unsealed court filings show OpenAI and Microsoft staff warned AI models replace news publishers
Internal records released in Manhattan federal court show Microsoft and OpenAI staff acknowledged their systems could replace original journalism, challenging the companies' fair-use defense in a copyright lawsuit brought by The New York Times and 11 other publishers.
Internal debate over data scraping
Unsealed court filings made public on 17 September 2026 in Manhattan federal court revealed internal discussions at Microsoft and OpenAI regarding the use of journalism to train artificial intelligence models. The documents were released as Judge Sidney H. Stein of the U.S. District Court for the Southern District of New York evaluates motions for summary judgment. The New York Times, joined by eleven other publishers including the New York Daily News, filed the infringement suit in late 2023, alleging unauthorized copying of millions of articles. In one internal 2023 document, Microsoft director of applied science and Northwestern University professor Brent Hecht characterized the mass ingestion of published work as an immense taking. He warned that scraping without consent or compensation could trigger a doom loop that harms model quality over time.
Millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft of unprecedented proportions
Substitutive impact on newsrooms
The legal filings cite OpenAI executives noting that chatbot systems function directly as replacements for traditional news publications. In a June 2023 memo, Nick Turley, who led the development team for ChatGPT, wrote that generative artificial intelligence posed an existential threat to media companies. In February 2024, Turley recorded that AI interfaces are largely substitutive and will grow increasingly substitutive as capabilities advance. During a deposition, Microsoft Chief Executive Satya Nadella stated that chatbots deliver information directly on the platform rather than requiring users to visit the primary source. Furthermore, a 2020 message from OpenAI co-founder and president Greg Brockman stated that the models showed high proficiency when completing sentences in New York Times articles.
Like whenever i have it complete in the middle of a sentence in a NYT article, it seems to complete the sentence on point.
Fair-use defense and corporate response
OpenAI and Microsoft maintain that training large language models on news archives constitutes fair use under United States copyright law. The companies argue that the technology transforms source texts into new outputs and does not compete unfairly with original newsrooms. Microsoft spokesman Alex Haurek stated that memos written by Hecht reflected personal views rather than official corporate policy. He stated that the company's legal filings explain why transformative use complies with copyright standards and why its Copilot assistant does not replace journalism. Steven Lieberman, an attorney representing the publishing groups, countered that the newly unsealed records show technology executives knew their systems were built to replace news services.
Microsoft's position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism
Legal timeline and industry disputes
The legal battle traces back to 2019, when Microsoft became an early financial backer of OpenAI and supplied computing resources for model training. OpenAI released ChatGPT in November 2022, prompting disputes across the publishing sector over unauthorized data ingestion. The New York Times and The Wall Street Journal have also engaged in a separate copyright dispute with search engine startup Perplexity. In June, Times publisher A.G. Sulzberger criticized technology companies for appropriating original reporting without compensation. The plaintiffs are asking the federal court to grant summary judgment finding OpenAI and Microsoft liable for unauthorized copying across each stage of their AI pipeline.
- Microsoft becomes an early investor in OpenAI and agrees to provide computing infrastructure.
- Greg Brockman notes models excel at predicting text from New York Times reporting.
- OpenAI publicly launches the ChatGPT conversational artificial intelligence tool.
- Nick Turley writes an internal OpenAI memo noting AI poses an existential threat to publishers.
- The New York Times files a federal copyright infringement lawsuit against OpenAI and Microsoft.
- Nick Turley records that AI products will become increasingly substitutive for news websites.
- Manhattan federal court unseals internal Microsoft and OpenAI communications in the lawsuit.


