The newly revealed court documents in the New York Times case against OpenAI and Microsoft are quite damning. The companies’ own documentation warned that they were starting a “fatal loop” that would damage the website, characterized their data extraction to train their models as the “largest theft of labor in the history of humanity” and that it made a “complete mockery of the idea of fair use.” Many of the most striking quotes in the document come from Microsoft’s Director of Applied Sciences, Brent Hecht. However, the company has attempted to distance itself from Hecht’s claims. Microsoft spokesperson Alex Haurek told The Verge that “These comments reflect an individual employee’s perspective, are not legal analysis, and do not represent the views of the company.” In a separate court filing, Jordan Usdan, general manager of data strategy and operations at Microsoft AI, characterized Hecht’s role as an adversary. He said Hecht “has divergent, academic, and forward-thinking views on how data ecosystems for AI should operate and is employed at Microsoft to contribute asymmetric, futuristic, and academic views…nor is he someone who speaks on behalf of Microsoft specifically regarding his theoretical views on the potential effect of AI on content creators.” But whether or not Microsoft wants to take responsibility for these comments, it’s clear that this came true. Google Zero is real! AI is eating the web! There are many more wild claims in the NYT filing from a variety of figures, including Satya Nadella, Sam Altman, and other OpenAI employees. Below are some highlights from the 92-page document. “An astonishing robbery.” The introduction quotes Hecht and OpenAI’s ChatGPT director (presumably Nick Turley) in a way that seems to show that the companies knew they posed an “existential threat” to publishers like the New York Times. Hecht calls ChatGPT and Copilot’s data collection the “largest theft of labor in human history” and says Microsoft’s defense makes a “complete mockery of the idea of ’fair use.'” Satya Nadella admits that chatbots have basically replaced search and eliminated the need to go directly to the source of information. But perhaps more damning is an internal Microsoft document that says: “Our AI content strategy has started a ‘doom loop’ that will harm the performance of our models and the entire web at the same time: it is highly unusual for an end product to threaten the economic fundamentals of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’” OpenAI co-founder Greg Brockman is more interested in the “trillions” of dollars he could make through commercial AI. Although Nadella was later quoted as saying, “anything that has a paywall has to be licensed,” an OpenAI representative admitted that he was “not aware” of any efforts to detect or remove paid content from training data. “Incredibly good at regurgitation” Internally, it seems that OpenAI was well aware of ChatGPT’s tendency to simply reproduce itself. copyrighted material “verbatim.” While acknowledging that “prevention of memorization” was important to “minimize copyright violations,” employees admitted that GPT-4 “memorized a ton of data and will therefore be incredibly good at regurgitation.” The presentation then cites several examples of ChatGPT generating long strings of text directly from articles in the Times, Mercury News, The Denver Post, LifeHacker and Eurogamer in response to queries. Microsoft knew how its widespread Internet scraping would be perceived, admitting that “almost no one intended to do it.” [sic] content they created to be used in this way, nor are they compensated for its use.” A “substitute for people’s work.” OpenAI’s policy director, Jack Clark, saw the writing on the wall, saying that it was “creating systems that replace the work of people who define the ‘culture’ of society.” Internal documents described ChatGPT as “the modern kiosk.” OpenAI’s Nick Turley is later quoted as saying that once you get a response from your chatbot, “no There is a good reason to click” on a link to the source. Destroying your own supply chain Microsoft is quoted as admitting that “LLMs are a product that destroys your own supply chain” because in many cases it is a substitute for your own training data. in this case they are perfectly consistent. He spoke of broad principles and ongoing changes in the way people find and consume information. Those observations should not be confused with conclusions on copyright issues before the Court, which Microsoft addresses in its filings.” But it seems pretty clear, based on this newly revealed document, that both Microsoft and OpenAI knew they were going to irreparably harm the publishing industry, the “millions of people” it employs, and, by extension, harm their own product, but they continued pursuing “millions” of dollars anyway—fatal loop be damned. Follow this story’s topics and authors to see more like this in your personalized feed. the home page and to receive email updates.Terrence O’BrienCloseTerrence O’BrienPosts from this author will be added to your daily email digest and homepage feed.FollowFollowSee all by Terrence O’BrienAICloseAIPosts from this topic will be added to your daily email digest and homepage feed.FollowFollowSee all AIGoogleCloseGooglePosts from this topic will be added to your daily email digest and your homepage feed.FollowFollowSee All GoogleLawCloseLawPosts in this topic will be added to your daily email digest and homepage feed.FollowFollowSee All LawMicrosoftCloseMicrosoftPosts in this topic will be added to your daily email digest and homepage feed.FollowFollowSee All MicrosoftOpenAICloseOpenAILPosts in this topic will be added to your daily digest Posts in this topic will be added to your daily email digest and homepage feed.FollowFollowSee All PolicyTechCloseTechPosts in this topic will be added to your daily email digest and homepage feed.FollowFollowSee All Tech