Friday, Jul 24, 2026 The claims desk. Receipts included. POWERED BY LENZ
IsThis

HEALTH

The Claim

AI chatbots, such as ChatGPT, provide medical advice that is consistently reliable and safe for users.

The Short Version

The claim that AI chatbots like ChatGPT provide "consistently reliable and safe" medical advice is not supported by the evidence. Multiple high-quality studies from 2024–2026 show ChatGPT gave incorrect advice in over 51% of medical emergencies, exhibited hallucination rates of 50–82%, and correctly identified conditions in fewer than 34.5% of real-world cases. ECRI designated AI chatbot misuse as the top health technology hazard for 2026. While chatbots show promise in narrow, controlled tasks, their performance is neither consistent nor safe for general medical advice.

Caveats

  • AI chatbots can hallucinate medical information with high apparent plausibility — studies document hallucination rates between 50% and 82%, meaning users may receive confidently stated but entirely fabricated guidance.
  • Strong performance on curated benchmarks or common scenarios does not translate to reliable real-world medical advice; studies of actual user interactions show dramatically lower accuracy rates.
  • General-purpose chatbots like ChatGPT are not regulated or validated as medical devices and should not be used as substitutes for professional medical consultation, especially in emergencies.

The Receipts

  1. AI chatbots and (mis)information in public health: impact on vulnerable communities - PMC

    PMC

  2. Role of Artificial Intelligence in Patient Safety Outcomes: Systematic Literature Review

    PMC

  3. Regulating AI in Medical Devices: FDA and EU Expectations by 2026 - PRP Compliance

    PRP Compliance

  4. Evaluation and mitigation of the limitations of large language models in clinical decision-making - PMC - NIH

    PMC - NIH

  5. AI chatbot misuse tops annual list of health technology hazards - RISE

    RISE

  6. Accuracy of Large Language Models When Answering Clinical Research Questions: Systematic Review and Network Meta-Analysis

    PubMed

  7. Comparative analysis of large language models in clinical diagnosis: performance evaluation across common and complex medical cases - PubMed

    PubMed

  8. On the limitations of large language models in clinical diagnosis | medRxiv

    medRxiv

  9. AI Chatbots Miss More Than Half of Medical Diagnoses, Study Finds - CNET

    CNET

  10. AI Chatbots Can Run With Medical Misinformation, Study Finds, Highlighting the Need for Stronger Safeguards | Mount Sinai

    Mount Sinai

+ 12 more sources — see the full list on Lenz

Filed Under

Artificial Intelligence ChatbotsChatGPTMedical Advice

More Fact Checks