New details have appeared in US publishers’ lawsuit against OpenAI and Microsoft. Partly unsealed filings quote Microsoft applied-research director Brent Hecht describing the large-scale copying of journalistic material as “the largest theft of labour in human history”. The filings also discuss a fear that chatbot answers could replace visits to news sites.

What the documents do — and do not — show

The case was brought by The New York Times, New York Daily News and other publishers. They argue that the companies copied protected articles for model training without permission and then built products capable of substituting for the original publications.

The new material is significant, but most of it appears in the publishers’ submission. Supporting exhibits remain sealed, so the quotations lack their full context. The court has not found that either OpenAI or Microsoft infringed copyright.

The question of scale

Publishers claim that intermediate training data contained tens of thousands of copies of their works and that a Common Crawl-based set contained millions of documents from nytimes.com. They also describe a Microsoft–OpenAI project known as Mango. Large counts alone do not decide the legal question.

The dispute turns on whether training is a transformative analysis of work or commercial copying that requires a licence, how the data was obtained, and whether the resulting product substitutes for the original on the market.

Why referral traffic matters

The filings cite an internal Microsoft statistic suggesting that links to The New York Times could receive much lower click-through rates in Copilot answers than in ordinary Bing results. The method has not been publicly disclosed, so it is not an independent measurement of all search traffic. Still, it captures the central economic concern: an AI answer may satisfy a reader before that reader reaches the publication that paid to produce the source material.

Bottom line

The unsealed documents make the arguments sharper, not the verdict clearer. They show that people inside the companies discussed the same tension publishers now put before the court: AI can summarise the web more conveniently, but that convenience can weaken the incentives that created the reporting in the first place. The court’s eventual answer will shape how training data is licensed and used well beyond this lawsuit.

Sources

  1. The New York Times — reporting on the filings
  2. TechCrunch — report on the unredacted filings
  3. The Washington Post — court-record reporting