Court documents reveal OpenAI and Microsoft scraped millions of news articles for AI training
World · 18 September 2026
Written by AI from multiple news reports
Court documents in a copyright case have revealed how OpenAI and Microsoft collected news content to train their AI systems. The documents were recently unsealed.
OpenAI scraped more than 10 million articles, and about one-third came from the New York Times. A joint project between the two companies collected copies of over 160,000 works from news publishers. OpenAI employees also found ways to get around paywalls and removed copyright notices from the material before using it.
One Microsoft manager called the practice possibly the largest theft of labour in human history. Both companies say their use of the content is legal because AI systems transform the material rather than copy it. The New York Times and other publishers disagree.