Internal concerns emerge over AI training and the future of journalism
Recently unsealed court documents have shed new light on internal concerns within Microsoft and OpenAI over the impact of artificial intelligence systems on journalism and the use of news content to train large language models.
The disclosures form part of a closely watched copyright dispute involving The New York Times, OpenAI and Microsoft. The newspaper and other publishers accuse the technology companies of using millions of copyrighted articles without authorization to train AI systems. OpenAI and Microsoft have defended their practices, arguing that the training process is transformative and protected under the US fair-use doctrine.
Among the documents made public were comments from Microsoft employees expressing concern that the large-scale use of journalistic material could ultimately weaken the very ecosystem on which AI systems depend. Microsoft applied science director Brent Hecht was quoted as warning that millions of people could view the large-scale collection of their work by AI companies as an unprecedented form of appropriation. Microsoft has stressed that such comments represented an individual employee’s perspective rather than the company’s official legal position.
Other evidence cited in the litigation concerns OpenAI’s internal discussions about the changing relationship between AI products and traditional news publishers. Nick Turley, who leads ChatGPT, reportedly described AI products as potentially posing a major threat to publishers and said their ability to substitute for existing sources of information could increase as the technology improves.
The documents also highlight concerns about the way AI-powered search and chat tools could alter the flow of traffic to news websites. As users increasingly obtain summaries and answers directly from AI systems, publishers have raised concerns that fewer people may visit the original articles, potentially affecting advertising, subscriptions and other sources of revenue.
Microsoft CEO Satya Nadella acknowledged during testimony that conversational AI can change how people obtain information, allowing users to receive answers directly through an AI platform instead of visiting the original website. Microsoft has maintained that such observations do not change its position on the copyright issues being considered by the court.
The legal dispute has become part of a much broader debate over the rules governing AI training. Publishers, authors and other copyright holders argue that technology companies should obtain permission or provide compensation when their works are used to develop commercial AI systems. Technology companies, meanwhile, contend that training models involves extracting statistical patterns from copyrighted material rather than reproducing the original works and that such use can qualify as fair use.
The disagreement remains unresolved in US courts. In September, OpenAI and Microsoft asked a federal judge in Manhattan to rule in their favor, while The New York Times, other news organizations and authors sought rulings supporting their copyright claims. The case could help determine how US copyright law applies to the training of generative AI systems and how the industry can use protected material to develop increasingly capable models.
The dispute is also expanding beyond the original plaintiffs. In September, The Seattle Times and Newsday filed a separate lawsuit accusing OpenAI and Microsoft of copying journalistic content without permission for AI training. The new case illustrates the growing number of legal challenges facing technology companies over the use of copyrighted material.
At the same time, the debate is moving toward a broader question about the future of the information ecosystem. AI developers need large quantities of high-quality material to improve their systems, while news organizations depend on sustainable revenue to continue producing original reporting. The legal and commercial arrangements established in the coming years could therefore influence how publishers license their archives, how AI companies obtain training data and how audiences access news online.
For now, the central issue remains contested. Publishers argue that unauthorized AI training can undermine the economic value of original journalism, while OpenAI and Microsoft maintain that their use of copyrighted material can be legally transformative. The courts will ultimately determine how existing copyright principles apply to these emerging technologies.
-
17:30
-
17:15
-
17:00
-
16:44
-
16:30
-
16:15
-
16:00
-
15:45
-
15:30
-
15:15
-
15:00
-
14:45
-
14:31
-
14:30
-
14:15
-
14:00
-
13:45
-
13:20
-
13:15
-
13:05
-
12:50
-
12:35
-
12:20
-
12:05
-
11:47
-
11:32
-
11:15
-
11:00
-
10:45
-
10:45
-
10:30
-
10:23
-
10:15
-
10:00
-
09:42
-
09:30
-
09:25
-
09:15
-
09:10
-
09:00
-
08:51
-
08:45
-
08:35
-
08:30
-
08:17
-
08:15
-
08:10
-
07:45
-
07:40
-
03:06
-
01:55
-
23:59
-
23:55
-
23:42
-
23:33
-
23:23
-
23:15
-
23:00
-
22:45
-
22:30
-
22:15
-
22:00
-
21:46
-
21:30
-
21:15
-
21:00
-
20:45
-
20:30
-
20:15
-
20:00
-
19:45
-
19:30
-
19:15
-
19:00
-
18:45
-
18:30