The authors aimed to explore the strengths and weaknesses of 2 different large language model (LLM) chatbots (ChatGPT4.0 and DeepSeek-RI) in providing reliable and access information about paediatric ophthalmology. The study used 44 multiple choice questions covering topics from the paediatric ophthalmology and strabismus section of the American Academy of Ophthalmology 2024–25 Basic and Clinical Science Course (17 strabismus, 6 neuro-ophthalmology, 4 retina/vitreous, 2 clinical optics and visual rehabilitation, 4 oculofacial and orbital surgery, 3 lens/cataract, 3 corneal/external disease/anterior segment, 2 general, 2 glaucoma, 3 intraocular tumours/uveitis and 1 trauma question). These were entered verbatim into LLMs. Model responses were evaluated by direct comparison to the official answer key without subjective interpretation or author interference. ChatGPT answered correctly for 82% (strabismus accuracy of 70% and 89% accuracy for the remainder. DeepSeek answered correctly for 93% (strabismus accuracy of 82% and 100% accuracy for the remainder. Although responses were not significantly different, DeepSeek outperformed ChatGPT for accurate responses. Accuracy issues were found to relate to insufficient training data.
Accuracy of AI chatbots in providing reliable paediatric ophthalmology responses
Reviewed by Lauren Hepworth
AI in pediatric ophthalmology: a comparative study of CHATGPT-4.0 and DeepSeek-RI performance.
CONTRIBUTOR
Lauren R Hepworth
University of Liverpool; Honorary Stroke Specialist Clinical Orthoptist, Northern Care Alliance NHS Foundation Trust; St Helen’s and Knowsley NHS Foundation Trust, UK.
View Full Profile
