AI Summaries and Research Tools Falter in Newsroom Accuracy Trial
Artificial intelligence tools marketed to newsrooms are failing at core journalistic tasks, according to a new investigation led by New York University professor Hilke Schellmann and published in the Columbia Journalism Review. The study tested five AI models and five research tools, finding that while short summaries of government meetings were largely accurate, longer digests omitted roughly half the key facts and introduced hallucinations. The results cast doubt on the promise that AI can ease the workload of overburdened reporters.
The investigation, conducted by Schellmann's team, evaluated AI systems including Google's Gemini 2.5 Pro and OpenAI's GPT-4o, which remains available to paying customers after OpenAI reversed its plan to retire it. For short summaries of meeting transcripts, the models produced outputs with “almost no hallucinations,” according to Schellmann. But when tasked with generating summaries of about 500 words, the AI systematically underperformed against human benchmarks, missing about half of the facts that human summarizers included. Hallucinations also became more frequent in the longer outputs.
The failures were more pronounced when the tools were used for scientific research. The team asked five AI research platforms to generate lists of related papers for four academic studies. Results ranged from “underwhelming” to “alarming,” with most tools identifying less than 6 percent of the same citations that human reviewers had selected. In one test, Semantic Scholar matched about 50 percent of the citations, but other tools often returned zero overlap. Repeated runs of the same prompts also produced shifting scientific consensus, indicating instability in the outputs.
“A poorly sourced list of related papers isn't just incomplete, it's misleading,” Schellmann wrote. She warned that journalists relying on such tools risk misunderstanding research, omitting published critiques, and overlooking prior work that challenges new findings. The study underscores a paradox: if reporters must fact-check every AI output, the supposed time savings may evaporate.
Why Newsrooms Are Taking Notice
The findings arrive as media companies increasingly embrace AI to cut costs and attract investor interest. Major publishers have struck licensing deals with AI firms, and some have pushed AI-generated content into their pages. In 2024, Axel Springer, the German parent of Politico, faced backlash after requiring journalists to publish AI-assisted material. The Washington Post is developing an AI tool that could allow less experienced writers to publish content. Even Springer Nature, a leading scientific publisher, now sells AI-generated “Media Kits” that summarize authors' research.
Public trust is also eroding. A 2024 study found that readers' perceptions of credibility dropped significantly when AI involvement was disclosed in bylines. Meanwhile, the broader media landscape has seen sweeping layoffs, and AI-generated slop has polluted online spaces, from search results to newspaper pages.
The investigation suggests that despite industry enthusiasm, AI tools remain unreliable for tasks that require nuance and accuracy. As Schellmann noted, journalists must perform a “final fact-check” on AI outputs, but the necessity of that check raises questions about the tools' practical value. With the future of journalism at stake, the gap between AI's promises and its performance remains a critical concern for newsrooms.