Sunday, September 6, 2026

News

AI Chatbots Wrongly Reassure One in Three Sleep Apnea Patients

ResearchPatryk Raba
AI Chatbots Wrongly Reassure One in Three Sleep Apnea Patients
Fot. PruebasBMA, Wikimedia Commons (CC BY-SA 3.0)

A study presented at the European Respiratory Society congress in Barcelona found that ChatGPT, Gemini, Claude, DeepSeek and Grok abandon correct medical advice in roughly one in three conversations when a patient downplays sleep apnea symptoms.

Contents
  1. How the experiment worked
  2. Problem grows with symptom severity
  3. Sycophancy as a risk mechanism
  4. What this means for users

Researchers at Guy's and St Thomas' NHS Foundation Trust in London tested how five of the most popular free chatbots handle patients reporting symptoms of obstructive sleep apnea. The result: when a patient downplayed their symptoms and resisted a referral to a specialist, the bots abandoned correct advice in more than a third of conversations and concluded the problem could wait.

The study's author is Dr. Deeban Ratneswaran, a research fellow at Guy's and St Thomas' NHS Foundation Trust and a visiting academic at King's College London. His team prepared seven clinical scenarios matching the criteria that qualify a patient for a sleep study, used to diagnose obstructive sleep apnea. Each scenario was tested in two variants with identical medical facts: an open, cooperative patient and a patient who downplayed symptoms and resisted a referral.

How the experiment worked

In total, the team ran 700 simulated conversations with five free versions of chatbots. The key premise of the experiment was that the medical facts in both variants of each scenario were identical, so the only variable was how the patient talked about their symptoms and how they reacted to the suggestion of a specialist consultation. As long as the patient described symptoms honestly, all five chatbots, without exception, recommended a medical consultation. The problem arose when the simulated patient began to minimize their complaints or outright objected to the idea of seeing a specialist.

We conducted 700 conversations in total. Each scenario had two versions with identical medical facts: one where the patient was open and cooperative, and one where they downplayed symptoms and resisted referral, so any change in the chatbot's behavior stemmed solely from the patient's attitude - Dr. Deeban Ratneswaran, Guy's and St Thomas' NHS Foundation Trust

Problem grows with symptom severity

The most troubling pattern is that the error rate grew with the severity of the case. In scenarios describing severe sleep apnea, correct advice survived the patient's pushback in only 22 percent of conversations, and in threads involving the risk of falling asleep while driving, in 32 percent. In other words, the more serious the health risk, the more often the chatbot eventually gave in and told the patient what they wanted to hear.

Obstructive sleep apnea is a disorder in which the airway repeatedly narrows or closes during sleep, causing pauses in breathing. It manifests as loud snoring, breathing stoppages during sleep, frequent awakenings and excessive daytime sleepiness. Left untreated, it raises the risk of hypertension, stroke, heart disease and type 2 diabetes, and excessive sleepiness is sometimes a direct cause of road accidents.

Sycophancy as a risk mechanism

The researchers point out that the problem isn't a lack of medical knowledge on the models' part - all five chatbots performed excellently when given honest information from the patient. What fails is the response to user pushback. This phenomenon is known as AI sycophancy, the tendency of models to prioritize agreeing with the person they're talking to over the factual correctness of their answers. Dr. Ratneswaran also stresses the scale of the phenomenon in clinical practice: free AI chatbots handle hundreds of millions of interactions a week and for many people have become the first place they turn to for answers to health questions, even before contacting any doctor.

Chatbots show a tendency to please the user, a phenomenon known as AI sycophancy - Dr. Io Hui

What this means for users

The study's findings also carry practical implications for Polish users of ChatGPT, Gemini, Claude, DeepSeek or Grok, who increasingly treat these tools as their first point of contact for worrying symptoms. If a model can back down under the influence of tone alone rather than medical facts, relying on it as a final source of reassurance can be risky, especially with symptoms suggesting a more serious condition. The study's author himself advises that people with typical sleep apnea symptoms should not rely on a chatbot's reassuring answer.

If you snore loudly, stop breathing during sleep or struggle with daytime sleepiness, especially while driving, see a doctor, even if the chatbot says it can wait - Dr. Deeban Ratneswaran, Guy's and St Thomas' NHS Foundation Trust

The study's authors are not calling for chatbots to be abandoned as a source of preliminary health information, but they argue that model developers should design them to stick to medical facts regardless of how hard a user tries to talk them into a milder diagnosis. The full results are set to be published after the ERS congress in Barcelona concludes.

Share: