Chatbots Flunk Memory and Drawing Tests, Raising Doubts on Medical Use
Leading artificial intelligence chatbots, including OpenAI's GPT-4o and Google's Gemini, showed signs of mild cognitive impairment when given a standard dementia screening test, according to a study published this week in The BMJ. The results, which challenge the notion that AI is ready to replace human doctors, found that the most capable model barely passed the threshold for normal cognition, while others scored well below it.
The study, conducted by researchers who administered the Montreal Cognitive Assessment (MoCA) to five prominent large language models, found that all of them struggled with tasks requiring visuospatial skills and executive function—abilities that are crucial for interpreting medical images and planning treatment. The findings are particularly relevant as hospitals and clinics increasingly pilot AI tools for diagnostic support.
How the Chatbots Scored
Among the tested models, OpenAI's GPT-4o achieved the highest score with 26 out of 30, just above the cutoff for normal cognition. Anthropic's Claude 3.5 Sonnet followed closely, while the two Google Gemini models (versions 1.0 and 1.5) scored the lowest at 16 out of 30, a result the researchers described as 'horrendous' in the study.
All the chatbots performed well on tasks involving naming, attention, language, and abstraction. However, every model failed at drawing tasks, such as connecting circled numbers in ascending order or sketching a clock to show a specific time—exercises that require spatial reasoning and planning. The Gemini models also failed a delayed recall test, where they had to remember a five-word sequence after a short interval.
The researchers noted that such deficits could undermine a chatbot's reliability in a clinical setting. 'These findings challenge the assumption that artificial intelligence will soon replace human doctors,' the study authors wrote, 'as the cognitive impairment evident in leading chatbots may affect their reliability in medical diagnostics and undermine patients' confidence.'
Empathy and the Anthropomorphism Trap
The study also flagged an 'alarming lack of empathy' in all chatbots, a trait that is a hallmark of frontotemporal dementia in humans. While the researchers acknowledged the fundamental differences between a biological brain and a language model, they argued that if tech companies market these AIs as conscious or human-like, it is fair to evaluate them using human cognitive standards.
This is not the first time AI has been scrutinized for potential cognitive gaps, but it is among the first to apply a validated clinical tool like the MoCA to chatbots. The study's authors emphasized that the goal was not to diagnose AI, but to counter a wave of research that overstates the technology's readiness for medical use.
The researchers concluded that neurologists are unlikely to be replaced by large language models anytime soon. Instead, they suggested, doctors may soon encounter 'new, virtual patients—artificial intelligence models presenting with cognitive impairment.'
For now, the study serves as a cautionary note for healthcare providers considering AI adoption, reminding them that even the most advanced chatbots have limitations that could affect patient care.