AI Scraping Theft: Microsoft’s Own Scientist Called It the Largest in History

Sep 21, 2026 | microsoft ai

In a nutshell

everything on the web starts with the domain

Big Tech has one public answer to the question of where its AI learned everything it knows: fair use. Perfectly legal, entirely transformative, nothing to see. But this week, the wall between what these companies say in court and what they said to each other came down a little. Newly unsealed filings in the New York Times' copyright case against OpenAI and Microsoft reveal that a senior Microsoft scientist, in private, did not call the practice fair use. He called it theft — and not a small one. He called it perhaps the largest theft of labor in human history. That phrase, from inside the house, is the story.

What the Filings Revealed

On 17 September, previously redacted material was made public in the copyright lawsuit the Times and other publishers brought against OpenAI and Microsoft back in 2023 — a case a federal judge allowed to proceed in 2025. According to the filings, Microsoft's own Director of Applied Science, Brent Hecht, wrote in an internal memo in early 2023 that the industry-wide practice of scraping the open web to train large language models was an astonishing theft of unprecedented proportions, and, in his words, perhaps the largest theft of labor in human history. He added that invoking fair use to justify it made, as he put it, a mockery of the concept. An OpenAI executive, separately, described the company's own models as posing an existential threat to the publishers whose work trained them. And one Microsoft document conceded the strangeness of the whole arrangement: that it is highly unusual for an end product to threaten the economic foundations of its essential suppliers, which is precisely what the AI business had done to its content supply chain.

The filings also describe the mechanics, and they are not flattering: paywalls bypassed without detection, training datasets assembled through mass scraping, and copyright notices deliberately stripped from the text before it was fed to the models. The scale is staggering — more than ninety thousand copies of Times-group works in a single mid-training dataset, and over two million documents drawn from the Times' domain alone in another. Microsoft's chief executive, Satya Nadella, testified that paywalled content should ideally be licensed, and that he would have required a retraining had he known it had been taken without permission.

The Gap Between the Podium and the Memo

Here is why this matters beyond one lawsuit, and it is worth stating carefully. The significance is not that AI models are trained on the open web — everyone has known that for years. The significance is the distance between the public position and the private one. In court and on stage, these companies argue, with real legal firepower, that training on copyrighted work is transformative fair use and harms no one. In their own memos, at least one senior scientist described the same practice in the language of theft, and executives worried in writing that they were gutting the very institutions they depend on. That gap is the thing. It is the difference between a company that believes its own defence and a company that mounts a defence it privately doubts. Whether or not the law ultimately calls this fair use — and that question is genuinely unresolved — the internal record suggests the people building these systems understood, early and clearly, exactly whose work they were taking and what it might do to them.

The Theft of Labor — the Phrase That Lands

Sit with the specific words the Microsoft scientist chose, because they are more precise than they first appear. Not theft of property. Theft of labor. Every article scraped represents hours of a human being's work — a reporter who made the calls, an author who spent the years, an editor who shaped the sentences. The models did not ingest abstract data; they ingested the accumulated labor of millions of people, none of whom were asked, and none of whom were paid. And this is the thread that runs through everything this site has argued about AI and human worth. We have written that the danger of AI was never the pretty pictures but the quiet devaluation of the people behind them. Here is that devaluation made concrete and admitted: the largest transfer of uncompensated human effort in history, its authors turned into raw material for a product that, by the companies' own internal reckoning, then goes on to compete with them and cut their traffic by as much as ninety percent. A doom loop, one document called it. The machine consumes the work, then starves the workers who made the work possible.

Claim and Counter-Claim

The case for the companies deserves a fair hearing, and it is not empty. Fair use is a real and important doctrine that has, historically, enabled search engines, libraries, and countless transformative technologies; training a model that learns patterns from text is arguably different in kind from republishing that text, and the law genuinely has not settled whether it crosses the line. Microsoft has formally disavowed its own employee's characterisation, stating that its court filings set out why these uses are lawful and why its products do not substitute for the journalism they learned from. An internal memo by one scientist, however striking, is a private opinion, not a legal finding — and much of what is now public comes from the plaintiffs' own brief, not from the still-sealed underlying exhibits.

The case against is that private candour tends to reveal what public argument conceals. When a company's own Director of Applied Science writes the word theft, unprompted and years before any verdict, and when its CEO concedes under oath that the material should have been licensed, the fair-use defence starts to look less like conviction and more like strategy. And the deepest objection is not even legal; it is moral and economic. Even if a court blesses the scraping as lawful, the arrangement that Microsoft's own document described remains real: an industry built on the uncompensated labor of the people it is now out-competing, hollowing out journalism and creative work as it scales. The honest synthesis: the legality is unsettled and may well favour the companies, and the internal record is nonetheless a damning portrait of an industry that knew the human cost of what it was doing, priced it at zero, and argued otherwise in public.

The European Perspective

For Europe, this American courtroom drama illuminates a fork the continent has, characteristically, already tried to take. The United States is fighting this out through the elastic, unpredictable doctrine of fair use — a single judge's ruling away from either vindicating or upending the entire industry. Europe chose a different instrument.

There is no broad fair-use defence in EU copyright law; instead there is a text-and-data-mining framework that lets rightsholders opt out of having their work mined, a neighbouring right that was created precisely to make platforms pay publishers for their journalism, and, now, transparency obligations under the AI Act that require providers to disclose the data their models were trained on. Where America asks a court, after the fact, whether the taking was fair, Europe tries to require, before the fact, that the taking be consented to and disclosed. Neither approach has yet forced a genuine reckoning — European publishers are as scraped and as squeezed as American ones, and enforcement lags far behind the technology.

But the unsealed memo is, in its way, the strongest argument yet for the European instinct. If even the companies' own scientists privately concede that mass scraping is an astonishing theft of human labor, then a legal regime built on requiring permission and payment, rather than on litigating fairness afterward, looks less like European over-caution and more like foresight. The deeper point is the one this whole affair forces into the open: the extraordinary machines now reshaping the world were built, in part, on work that millions of people did, that no one paid for, and that the builders themselves knew, in private, they were taking.

A civilisation that wants both powerful AI and a living culture of human creation has to decide whether that is a foundation it can live with. Europe, at least, is trying to write down that it cannot.

We are not first. We are right.