AI Health Chatbots Fall Short: Study Finds Misdiagnoses
AI Health Chatbots Fall Short: Study Finds Misdiagnoses
A UK study finds AI health chatbots misidentify conditions and offer wrong guidance, performing no better than basic web searches in real scenarios.
If you’re tempted to turn to Dr ChatGPT the next time you feel unwell, think again. A UK study tested how well chatbots help people pinpoint health problems and decide when to seek care, using nearly 1,300 participants and 10 different scenarios. Participants were asked to describe the problem and outline the next steps after interacting with one of three chatbots or with standard internet searches. The chatbots tested were GPT-4o, Llama 3, and Command R+. A control group used internet search engines for comparison.
Results showed the chatbots identified the underlying health problem only about a third of the time, while correctly advising on the next course of action in roughly 45% of cases. In other words, the performance of the AI tools was no better than the control group relying on search engines. The researchers note a striking gap between AI benchmarks and real-world usefulness, with chatbots performing well on medical exams but faltering in everyday patient conversations.
A key reason, they say, is a communication breakdown: real patients often don’t provide all the relevant information, and chatbots may miss warning signs or urgent needs because they lack the context that a clinician would gather in a visit. Rebecca Payne, a co-author from Oxford University, said, “Despite all the hype, AI isn’t ready to take on the role of the physician. Patients need to be aware that asking a large language model about their symptoms can be dangerous, giving wrong diagnoses and failing to recognise when urgent help is needed.”
The study underscores the importance of not treating AI chatbots as medical triage tools. While AI can be a helpful supplement for information, relying on them for diagnosis or treatment decisions could lead to delays in recognizing serious conditions or pursuing inappropriate actions. As AI continues to improve, researchers emphasize the need for clearer guidance on how and when to seek professional medical advice, and for better ways to collect complete information during virtual interactions.
Overall, the findings suggest that the current generation of AI health chatbots should be used with caution and literacy rather than as a substitute for professional medical care.