News

Legal Filing: OpenAI and Microsoft Knew They Were ‘Stealing’ AI Data

4 min read Editorial

A new legal filing has reignited the bitter OpenAI v. Anthropic battle with a striking new accusation in the OpenAI Microsoft legal filing: that executives at both companies knew their AI training data was “stealing,” yet pushed forward regardless. According to the filing, as covered by Thurrott, the parties allegedly prioritized a “mad bid” to advance their models over the questions of where that data came from.

What the filing alleges

The core of the document is a pointed attack on how the two companies approached AI training. The filing argues that the people running OpenAI and Microsoft were fully aware of the data-harvesting practices at play, and that they chose to proceed anyway. The language is deliberately forceful, framing the training of generative models not as a routine engineering decision, but as a deliberate crossing of a line.

For readers who haven’t been following the case, the word “stealing” is doing a lot of work. In litigation, accusations like this are weapons, not neutral descriptions. What one side calls “stealing,” the other may call “lawful data collection” or “fair use.” The filing is one party’s argument, not a court’s verdict.

Advertisement
A close-up of a thick legal binder open to a page of dense terms, a steel nib pen resting on top, soft office lighting,
The filing's forceful language is one side's argument, not a court verdict.

The OpenAI v. Anthropic backdrop

This filing doesn’t exist in a vacuum. It’s part of the OpenAI v. Anthropic lawsuit, filed in early 2025, that has dominated AI news for months. OpenAI claims that Anthropic’s Claude models were built using scraped copies of a proprietary model it calls GPT-2.5. OpenAI says Anthropic “stole” that data; Anthropic has denied the charge and argued the case is thinner than it appears.

What makes this case notable is the technical detail. Unlike typical trade-secret suits, the dispute hinges on whether specific model outputs resemble a specific model’s training data. Experts have publicly analyzed the evidence, and the back-and-forth has been unusually public. That context matters: the new filing’s accusation that OpenAI and Microsoft were the ones “stealing” is, in part, a counterattack in this same war.

Why Microsoft is named

Microsoft’s inclusion here isn’t a coincidence. The software giant is OpenAI’s primary backer, having invested billions and holding a stake in the startup. Microsoft also runs much of the infrastructure that powers OpenAI’s models on Azure cloud servers. Naming Microsoft pulls the deeper wallet and the data-center operator into the story.

For IT professionals and enterprise users, that linkage is worth noting. Many organizations run OpenAI-powered tools through Microsoft 365 Copilot and Azure. If the question of where that training data came from becomes a legal or compliance issue, the trail leads back to a company millions of businesses already depend on.

Microsoft Azure-style cloud server racks glowing blue in a dark data center, with a small OpenAI logo-shaped light panel
Microsoft runs much of the infrastructure powering OpenAI's models on Azure cloud servers.

What this means for you

On the surface, this looks like a corporate legal fight. But the underlying question—how AI models are trained on data scraped from the web—touches every user who’s ever typed something into a chatbot. If a court finds that a company “knew” the data was questionable and proceeded, it could set a precedent that ripples through the whole industry.

For the everyday user, the practical takeaway is simpler. Right now, there’s no recall, no build, no patch to install. This is a courtroom matter, not a Windows update. But it may shape what tools you’re allowed to use at work, and how your data is treated by the AI services you rely on.

How to stay informed

This is a developing story, and the specifics are still emerging. The filing is one side’s argument; there will be responses, rebuttals, and possibly a judge’s ruling. Until then, treat the forceful language with caution. For verification, follow the coverage on Thurrott, and watch official statements from OpenAI, Anthropic, and their legal teams. As with any AI-adjacent announcement, the companies’ own official positions should always be treated as the authoritative source.

Source: Thurrott.com

Over to you: As the OpenAI v. Anthropic case unfolds, which side do you trust to have the truth about how AI models are trained?

Advertisement
Share:
Editorial
Written by
Editorial

Windows & Microsoft news editor at 9to5Windows. Covering everything from Windows 11 builds to enterprise updates.

Advertisement