Share This

 

Patients are beginning to turn to large language model chatbots for healthcare information. As clinicians, we are not only required to know when they are wrong, but why they are wrong – and what we can do about it.

 

Recently, I evaluated artificial intelligence (AI)-generated patient information on Descemet membrane endothelial keratoplasty (DMEK) surgery [1]. I predicted running into some inaccuracies, but what I did not predict was how confident and specific these mistakes would be.

 

 

One model instructed patients to keep their heads elevated during recovery. This is not correct and plausible enough in its delivery for patients to follow without question. Another stated that vision would recover to baseline within one to two months, but in reality, can take three to six months. What I found most striking was a chatbot informing patients that they would undergo laser peripheral iridotomy (PI) two weeks before their procedure. This practice has been challenged since 2018 [2], as it can also be performed intraoperatively instead or omitted altogether, with studies showing 'PI-less' DMEK having comparable outcomes for visual acuity and graft rejection [3]. What makes me wary is that all of these errors were delivered with the same authority and confidence as the accurate information surrounding them.

This is not unique to a particular chatbot, as no model could produce the quality of the clinician-written leaflet, with several containing misinformation as mentioned above [1]. This begs the broader question: what happens when our patients come across this misinformation and what should we, as clinicians, be doing about it?

The anxious patient at home

Imagine a patient a few days after their DMEK surgery. They feel anxiety around whether their recovery is progressing normally with a follow-up appointment not for another week. They turn to a chatbot that gives them an easy-to-read, professional-sounding summary. It offers reassurance that full recovery will take one to two months and advice to keep their head elevated. To them, it reads like the leaflets provided at the hospital or sounds like something doctors might say. But it is wrong. The patient has no means of knowing that.

This scenario is fictional but not far-fetched. A major study conducted by King's Health Partners, responsible for the Policy Institute at King's College London, found that one in seven UK adults are turning to AI rather than their GP [4]. We know patients use Google for health advice but with the rise of AI, patients are getting something traditional search engines could never offer. Conversational responses from chatbots that feel personable, that feel trust-worthy. Delivering information to patients in a tone and structure that mimics clinicians.

Why AI gets it wrong

To appreciate why this matters, we need to understand how misinformation arises. Large language models do not have knowledge like clinicians do. They use patterns learned from large amounts of data scraped from the internet to generate a body of text, predicting the most statistically probable word after another [5] and fine-tune this with human feedback to sound conversational [6]. The result is an output that is often fluent and impressively accurate, with no mechanism to distinguish correct fact from plausible misinformation.

This is the phenomenon you have probably heard of: ‘hallucinations’, which leads to potentially dangerous errors [6]. For the laser iridotomy error mentioned earlier, we know the chatbot did not invent this as a concept but likely encountered this in out-dated literature. Not being able to separate past and present practice, it delivered this as up-to-date advice. This makes it harder for patients to recognise the information as wrong.

Through this same mechanism of using vast amounts of information gathered from the internet, chatbots can draw on information from multiple procedures that are closely related. An example here would be DMEK and Descemet stripping endothelial keratoplasty, which are both corneal transplant techniques. Although similar in theory, they are different in terms of postoperative instructions and recovery timelines. As a result, chatbots can produce advice that is an amalgamation of multiple procedures rather than just one specifically, making advice plausible but not quite right.

We see that chatbots treat everything with the same confidence, unable to weigh the significance or understand the repercussions associated with statements they make. For example, they do not stress the importance of postoperative positioning in DMEK being critical to the recovery process, or recognise they could cause unnecessary distress to patients through incorrect recovery timelines. All information is uniformly assured, whether statements are correct or not. The absence of nuance and caveats seem to be the most striking difference between a clinician’s communication and AI-generated advice. This is fundamental to honest, safe communication and is something chatbots cannot replicate as it stands.

An old problem in a different font

Chatbots and AI are new, but misinformation is certainly not. Patients have been using the internet to self-diagnose for over two decades now, coining phrases such as ‘Dr Google’, often with worries of alarming differentials and bleak prognoses [7]. We have learnt to recognise unreliable information through skills we have all been taught, even at school. This includes, recognising unverified sources such as forum sites and Wikipedia or identifying lack of authorship. Applying these skills on AI-generated information is more difficult as responses have been designed to sound like they are reliable and trustworthy – they echo many of these sources. They just cannot differentiate between the good and the bad ones when producing outputs.

Ophthalmology being relatively niche means less patient-facing information available on the internet compared to more generalised specialties such as cardiology. Also, the procedures are more specialised than others. Due to this, the training data chatbots use to produce ophthalmology-related outputs are limited. Therefore, the scope for error is greater.

Giving credit, where credit is due

This does not mean to say that chatbots and AI do not hold value today. At the time of our study on AI-generated patient information leaflets for DMEK , Claude 3.7 Sonnet achieved 77.8% on readability, reliability and accuracy [1]. Bearing in mind, the NHS leaflet scored 92% [8]. This is genuinely impressive for such a new technology, and we can already see rapid improvement in chatbots with every new generation.

A very plausible future could include chatbots producing first drafts of patient information for clinicians to review and verify. It could adjust to the level of literacy depending on the patient, making healthcare information more accessible. It could produce literature on all procedures and treatments where otherwise would not have been possible due to constraints in time and resources. We know AI will have a role in patient education in the future, and so it’s vital that we follow this closely and with clinical oversight. Despite being encouraging, we know AI is not immune to errors and therefore, for patients’ safety, clinician involvement remains essential.

What can we do as clinicians?

It’s unrealistic in my opinion to tell patients to avoid using AI at all costs. A more realistic method would be more robust. The most important thing we can do is to be aware. If working in a subspeciality or performing specific procedures, we must consider testing the outputs of different chatbots on these topics. My own experience doing this with the DMEK procedure was eye-opening and I suspect other clinicians would be surprised with what claims they find in their own specialities. This allows us to understand the landscape of misinformation and provide the first step in navigating it.

When patients come with specific expectations on treatment or recovery, it’s always a good idea to ask where they have been reading. This isn’t dissimilar to what we already do with information patients bring from Google. Directing them to verified sources such as the NHS website remains the gold standard for patient information. This can reduce the chances of misinformation getting to patients from chatbots.

Avoid dismissing AI but rather engage with it because AI in healthcare is not going away. Understanding how this technology works and recognising its pitfalls allows us to draw safe boundaries with AI and support patients already using these tools to ensure a safer implementation of this rapidly evolving technology.

 

TAKE HOME MESSAGES
  • Chatbots can produce patient information that is confidently written but can be inaccurate, often in the form of outdated advice presented as current clinical practice.
  • Misinformation arises from ‘hallucinations’, limited training data, blending of information across similar procedures and the inability to weigh clinical significance.
  • AI-generated patient information is more difficult to identify as unreliable compared to traditional sources due to it mimicking the tonality of professional clinical literature.
  • AI is improving rapidly and has potential in patient education, but no model matches clinician-generated information as it stands.
  • Clinicians should continue sign-posting patients to verified sources and become familiar with AI outputs relevant to their specialty. 

 

 

References

1. Massraf B, Chan KKC, Jain N, Panthagani J. Assessing accuracy, readability & reliability of AI-generated patient leaflets on Descemet membrane endothelial keratoplasty. Eur J Ophthalmol 2025;36(1):5–12.
2. Livny E, Bahar I, Levy I, . "PI-less DMEK": results of Descemet's membrane endothelial keratoplasty (DMEK) without a peripheral iridotomy. Eye (Lond) 2019;33(4):653–8.
3. Arora A, Sahu SK, Barik U, et al. Outcomes of DMEK without peripheral iridotomy and a review of the literature. Indian J Ophthalmol 2025;73(8):1197–201.
4. https://www.kcl.ac.uk/news/one-in-seven-people
-have-used-ai-instead-of-seeing-a-gp-study-finds

[Link last accessed August 2026]
5. Idan D, Einav S. Primer on large language models: an educational overview for intensivists. Crit Care 2025;29(1):238.
6. Kim TW. Application of artificial intelligence chatbots, including ChatGPT, in education, scholarly work, programming, and content generation and its prospects: a narrative review. J Educ Eval Health Prof 2023;20:38.
7. Goldberg I. Dr. Google will see you now: But will he make you sick? Taiwan J Ophthalmol 2024;14(3):371–5.
8. https://www.royalfree.nhs.uk/patients-and-visitors/
patient-information-leaflets/having-a-dmek
-or-dsaek-corneal-transplant

 

Declaration of competing interests: None declared.

 

Share This
CONTRIBUTOR
Bnar Massraf

Peterborough City Hospital, UK.

View Full Profile