Share This

This study evaluated the quality and readability of ChatGPT-generated patient educational materials for 4 common paediatric ophthalmology procedures, in both English and Spanish languages. Responses were compared to the American Association for Pediatric Ophthalmology and Strabismus (AAPOS) patient handouts. The 4 procedures were (1) strabismus surgery without adjustable sutures, (2) strabismus surgery with adjustable sutures, (3) paediatric cataract surgery with or without placement of intraocular lens, and (4) nasolacrimal duct probing with or without intubation. ChatGPT was asked to produce a 1000-word explanation using patient-accessible language at 6th grade level. Responses were scored for quality (assessed by the quality of generated language outputs for patients tool) and readability (assessed by the Flesch-Kincaid grade level, Gunning fog score and SMOG index). English responses were scored by 3 native English-speaking paediatric ophthalmologists and Spanish responses by 2 native Spanish-speaking paediatric ophthalmologists. For accuracy, all AAPOS materials received maximum scores. For ChatGPT, mean scores were from 2.36 to 3.33 (mean 2.79 ±0.79); significantly worse than AAPOS materials. For readability, the mean grade level of ChatGPT was 8.6 ±1.88. Overall, ChatGPT generated materials that were similar in readability, but inferior in quality, to AAPOS. English and Spanish responses and scores were similar. ChatGPT provided several instances of misleading information which could adversely affect patient expectations and outcomes. Caution is needed when generating such materials.

Quality and readability of patient educational materials generated by ChatGPR-4o for pediatric ophthalmologic surgeries. 
Yang A, Reid M, Nguyen A, et al.
JOURNAL OF PEDIATRIC OPHTHALMOLOGY AND STRABISMUS
2025;62(6):421–7.
Share This
CONTRIBUTOR
Fiona Rowe (Prof)

Institute of Population Health, University of Liverpool, UK.

View Full Profile