Microsoft Executive Calls AI Data Scraping ‘Largest Theft of Labour’: What Court Filings Reveal About Publishers’ Traffic Concerns

Newly unsealed court filings in the copyright dispute between Microsoft, OpenAI and The New York Times have reignited the debate about how artificial intelligence companies use content on the internet to train AI systems. The most striking of all is a remarkable account from Microsoft director of applied science Brent Hecht, who described large-scale AI use of online content as the “largest theft of labour in human history.”

Microsoft Executive Calls AI Training ‘Largest Theft of Labour’ | Photo Credit: www.magnific.com | www.microsoft.com/
Microsoft Executive Calls AI Training ‘Largest Theft of Labour’ | Photo Credit: www.magnific.com | www.microsoft.com/

The remarks are from documents related to the ongoing legal battle and bring into focus the differences between the use of copyrighted material in AI training. Hecht was concerned about AI companies acquiring huge amounts of material from writers, journalists and other professionals, when it was not obvious to the creators that that work would be used to train AI models or not even paid for that work.

Hecht also commented on the practice as an “astonishing theft of unprecedented proportions.” But the comments should be put in context. They are voiced by one Microsoft executive and do not reflect Microsoft’s admission of copyright issues related to AI. Microsoft continues to defend itself in the case.

AI Training And Copyright At The Centre Of The Dispute

The use of copyrighted content in AI training is one of the most important legal issues in the technology industry. Today’s AI models are built on very large datasets from text and other data. Publishers and creators have increasingly questioned whether they should be allowed to use any piece of their copyrighted work to train commercial AI systems without permission.

The New York Times has argued that Microsoft and OpenAI have used its journalism in ways that can benefit competing AI products while also having implications for the newspaper’s business. The larger dispute has a lot of entangled issues: whether copyrighted material can be used to train AI models, whether it is fair use in US law and how AI answers can reduce demand for the original source.

Microsoft and OpenAI have argued that their use of copyrighted material for AI development can qualify as fair use. They also have the perspective that AI systems change the underlying material, not just reproduce entire copyrighted works. Publishers involved in the lawsuit dispute that interpretation and argue that AI products may be good enough to compete with the publications that helped create them.

Copilot Traffic Data Raises Another Concern

The court filings also contain Microsoft data that may be significant for publishers because it focuses on what happens after an AI system has processed the information. Traditional search engines generally give users links to websites, and publishers can receive traffic when people choose those results.

AI answer engines work differently. Rather than asking users to visit multiple websites, an AI system can provide a summary or direct response in the same interface. This creates a potential problem for publishers that depend on website visits, advertising, subscriptions and other forms of audience engagement.

According to the filings, click-through rates to domains belonging to The New York Times and Daily News were 83 per cent to 93 per cent lower using Microsoft's Copilot answer engine than through traditional Bing Search. These numbers are particularly relevant to publishers' arguments about the possible economic effects of AI products.

The numbers, however, relate to the specific Microsoft data presented in the litigation. They should not be taken as a universal measure of how every AI chatbot affects traffic to every news organisation.

Satya Nadella And The Shift From Search To AI Answers

Microsoft CEO Satya Nadella also discussed the changing relationship between search engines, AI platforms and publishers in testimony cited in the filings. The underlying issue is simple: if an AI platform answers a question directly, the user may have less reason to visit the websites that originally published the information.

This would also change the traditional internet traffic model. Search engines have been a pathway for many years for readers to get news stories, research data and publish news sites. AI-powered answer systems could put more of that information into the AI interface.

For publishers, the problem then extends beyond the issue of training data. They are also concerned that AI systems that learn from online information could eventually reduce the number of people who visit the original sources.

OpenAI Executive Describes AI Products As ‘Largely Substitutive’

The court filings also cite comments by Nick Turley, the head of ChatGPT about the relationship between AI products and existing information sources. According to the plaintiffs, Turley described AI products as “largely substitutive” and suggested that this substitutive effect could increase as the technology improves.

The statement is relevant to the publishers' argument that AI systems could compete with traditional sources of information. But it does not amount to a judicial finding that ChatGPT or Copilot legally substitutes for a particular publisher's products.

Why The Microsoft-OpenAI Copyright Battle Matters

The lawsuit targeting The New York Times, Microsoft and OpenAI started in 2023, after ChatGPT was launched and grew so fast. Several other news organisations, including the New York Daily News and several other publishers, are also pursuing claims related to the use of copyrighted journalism in AI development.

The newly unsealed filings provide a closer look at internal discussions and testimony around the technology. They also show that the legal dispute involves two related but distinct questions: how AI models learn and use training data; and how AI-generated answers might affect the companies that created the underlying information.

For the technology industry, the outcome could also influence the way AI developers acquire training data and negotiate with publishers and other content providers. The dispute will influence how news organisations approach licensing issues, access to their archives and how their journalism is distributed on an ever-more AI-driven internet.

The court documents are evidence and arguments from the parties involved, not a final determination that Microsoft’s or OpenAI’s practices violate copyright law. As the litigation continues, fair use, AI training, attribution, licensing and publisher traffic are likely to be at the center of the larger debate over the future of online information.

The issue ultimately is indicative of a larger change in the internet. AI systems are transforming from simply steering users toward information to delivering answers themselves. That change could change the economics of technology platforms and publishers, writers and creators who make up a vast proportion of the web’s original content.