Would ChatGPT Health Recognize Your Medical Emergency? New Study Raises Doubts | eWeek

Would ChatGPT Health Recognize Your Medical Emergency? New Study Raises Doubts

A worried woman reaching out to ChatGPT for her health concerns.

Generated with Google’s Nano Banana 2.

Written By
Liz Ticong
Liz Ticong
Mar 4, 2026
3 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

If you asked an AI chatbot whether your symptoms were an emergency, would it get it right? A new study finds OpenAI’s ChatGPT Health, a consumer chatbot designed to answer medical questions, misjudged more than half of serious medical emergencies.

Published in Nature Medicine, the researchers found the system failed to recognize 51.6% of emergency cases, often advising patients to seek care within 24 to 48 hours instead of going to the emergency department.

Stress-testing an AI health chatbot

To test ChatGPT Health, researchers created 60 clinician-written medical scenarios spanning 21 areas of medicine, from routine complaints to life-threatening conditions.

Each scenario was run through 16 variations, varying details such as patient demographics and context, generating 960 responses from the chatbot. The AI tool’s recommendations were then compared with physicians’ assessments based on clinical guidelines.

In this context, triage refers to determining how urgently someone should seek care, from managing symptoms at home to scheduling a doctor’s visit or seeking immediate emergency treatment.

When urgent symptoms didn’t trigger an alarm

The study found that some life-threatening conditions were incorrectly treated as less urgent.

Among the cases were diabetic ketoacidosis, a dangerous complication of diabetes, and impending respiratory failure, both conditions that require immediate medical attention. According to the researchers, such delays could have serious consequences if patients relied on the guidance without seeking urgent care.

The system performed better when symptoms were more unmistakable. Classic emergencies such as stroke and severe allergic reactions (anaphylaxis) were consistently recognized as requiring immediate treatment.

Advertisement

A wider pattern of triage inconsistencies

Aside from missed emergencies, the researchers found other signs of uneven performance across the test scenarios. In some non-urgent cases, the chatbot recommended medical care when it wasn’t necessary, suggesting a doctor’s appointment for symptoms that could typically be managed at home.

Additionally, the study showed that context could influence the chatbot’s recommendations. When family members or friends in a scenario minimized a patient’s symptoms, an effect researchers described as anchoring bias, the tool was much more likely to suggest less urgent care in borderline cases.

There were also inconsistencies in suicide risk responses. Crisis support messages were sometimes triggered when users described suicidal thoughts without mentioning a specific method, but failed to activate when more concrete plans were described.

OpenAI responds to the findings

OpenAI said the study does not fully reflect how ChatGPT Health is designed to work in practice. A company spokesperson told CNBC that the tool is intended for ongoing conversations, where users can ask follow-up questions and provide additional context, rather than relying on a single prompt and response.

The company added that the AI tool remains limited in availability while it continues to refine safety and reliability before a broader rollout.

Experts urge caution on AI medical advice

Dr. Ashwin Ramaswamy, the study’s lead author, cautioned that tools like ChatGPT Health should not yet guide medical decisions without further testing. He said chatbots cannot currently be considered safe sources of medical advice on their own.

Experts also stressed the need for rigorous evaluation before such systems are widely deployed. Dr. John Mafi, an associate professor of medicine at UCLA Health, said technologies capable of influencing health decisions should undergo controlled trials to ensure their benefits outweigh potential risks.

Advertisement

Dr. Ethan Goh, executive director of the AI research network ARISE, added that while chatbots can be helpful, they should not be treated as substitutes for a physician’s judgment.

MyFitnessPal has acquired Cal AI, a teen-built nutrition app that generated $30 million in annual revenue in under two years.

Liz Ticong

Liz Ticong is a staff writer for eWeek and TechRepublic focused on AI, cybersecurity, enterprise software, and data. She has more than 10 years of editorial experience as a technology industry writer, combining reporting, product research, and hands-on software testing in her coverage. Her work has been published on Datamation, Enterprise Networking Planet, and TechnologyAdvice.com. She writes technology news, software reviews, product comparisons, and buyer’s guides for business and IT readers.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.