Much is said about the potential of artificial intelligence (AI) to transform clinical practice, particularly with the advent of foundation models [1]. The newest models are ‘pretrained’ on text, image, tabular and other forms of data to produce architectures that can then be ‘fine-tuned’ to complete a very broad range of tasks with close to or exceeding human expert performance levels [1].
Even before foundation models and generative AI emerged, impressive performance of glaucoma-related deep-learning algorithms for diagnosis and prediction of progression suggested that clinicians may be assisted by AI applications in the near future [2].
However, despite thousands of AI models being developed, validated and published in recent years, glaucoma clinics look very similar today to a decade ago. Glaucoma remains the more frequent cause of irreversible blindness, and its prevalence is expected to rise [3]. Glaucoma continues to be the greatest single source of strain on outpatient resources due to large numbers of patients with confirmed or suspected disease requiring monitoring and treatment titration [4]. Leveraging new technology to mitigate the burden of glaucoma for patients and clinics should be a high priority.

Figure 1: The Gartner hype cycle. Courtesy of Jeremy Kemp (CC BY 3.0).
The situation is consistent with the early stages of the Gartner hype cycle (Figure 1): Initial excitement with the emergence of new technology, followed by disillusionment as changes in day-to-day work fail to live up to the predictions based on early proof-of-concept work. For AI to improve glaucoma care, a shift in outlook is needed: emphasis on applications instead of technology, better directed validation work and clearer frameworks that allow clinicians and patients to take advantage of useful tools.
Thinking in terms of applications rather than technology
For researchers, using available AI technology with available data is the most convenient way to generate results that may be informative, but this rarely results in tools that can be used in clinic. It is much more challenging to develop an application to address a specific issue, but these tools are much more likely to be deployed and used.
Thinking about applications is also more natural to clinicians who best understand the difficulties and pain points encountered in their work. By deciding on a target application, it may also be possible to select technologies to maximise efficiency or minimise regulatory barriers, rather than being beholden to the newest generative AI model. Frequently, simpler computational approaches work out best in terms of accuracy, cost and implementation [5]!

Figure 2: Applications of AI in glaucoma, targeted towards pre-clinic decisions (A), clinician assistance in clinic (B), and using information from clinic to make decisions (C). Courtesy of Arun Thirunavukarasu et al (CC BY 4.0) [5].
The patient journey provides a useful framework for considering applications of glaucoma AI (Figure 2) [6]:
- Before clinic, scarce appointments require careful triage and prioritisation to ensure that patients at risk of vision loss are seen promptly, while patients at little to no risk are safely discharged.
- In clinic, AI tools can assist clinicians by flagging abnormalities (e.g. suspicious optic disc features) or providing predicted diagnoses or risk of disease progression.
- AI-enhanced decision-making could ensure that follow-up intervals or offers of medical, laser or surgical treatment are all aligned with available evidence and gold-standard practice.
Clinical studies of glaucoma AI interventions provide valuable data on the proposed use case but also show how applications may be adjusted to provide value. For instance, asymptomatic screening using AI-analysed fundus photographs and intraocular pressure (IOP) in the primary care setting has shown potential, but cost-effectiveness is limited by the number of false positive cases referred for work up [7]. The false positive rate drops as the prevalence of glaucoma rises in the tested population, and so targeted screening (such as through polygenic risk scores or even simple self-reported health data) and model use to triage community referrals to glaucoma services may prove more useful [7-9].
In diagnosis and decision-making, AI systems have shown superior performance to specialists tasked with interpreting fundus photographs and optical coherence tomography (OCT), as well as in providing advice or answering questions when provided with clinical information [10,11]. However, managing glaucoma is complicated: untested aspects include history-taking and consideration of patient-specific factors (e.g. treatment adherence, appropriateness of procedures). Clinicians are still needed to collect information and counsel patients, and so the challenge is now how to incorporate high-performance AI to improve decision-making and outcomes. Options range from passive presentation, such as by augmenting existing investigation printouts, to highlighting abnormalities or information that may prompt changes to management, and even more active suggestions of diagnoses or treatment plans.
Improving AI validation will inspire confidence in new tools
Unfortunately, a large proportion of the AI validation evidence base is uninformative. Incomplete reporting, internal validation with highly curated datasets, and observational study design all make drawing generalisable conclusions very difficult [2,12]. This is in part due to researchers fixating on technology – drawing on metrics and methodology from computer science – rather than applications and how best to test them in clinical research studies. Designing clinical research studies to show whether an application should be used is more challenging: ethics applications, patient recruitment and follow-up with outcomes capturing benefit all require funding and time while carrying risks of ‘failure’ which may put off developers. Moreover, these studies presuppose the development of applications that can improve outcomes significantly, which is itself a very difficult undertaking.
There are many initiatives that aim to improve the standard of AI validation methodology and reporting. Reporting guidelines aim to encourage best practices, with examples including the Chatbot Assessment Reporting Tool (CHART) [13]. CHART aims specifically to ensure reporting is transparent even as generative AI technology continues to evolve rapidly, which has been a challenge for previous guidelines. Regulators will also shape the validation landscape through the bar they set in terms of study design and results before granting approval to new applications.
However, it is crucial that researchers are not constrained by overly prescriptive guidance. Applications with variable functions and risks must be evaluated differently. Assistive tools with little direct impact on clinical care cannot reasonably be tested in randomised-control trials with clinical outcomes. Autonomous tools taking on the responsibility of a clinician should in contrast undergo pragmatic trials and post-deployment monitoring to ensure patients are safeguarded from harm. In general, experimental studies – where a new application is tested against usual care rather than looking at ‘before-after’ effects without a control arm – are preferred, but these must be tailored to the nature of the AI application [14].
Overcoming validation and governance barriers
Technology-based regulation is inherently controversial. Limitations based on computational power, model architecture or recency risk being rendered obsolete as development accelerates or developers find workarounds. Application-directed regulation is more appropriate, and fits into existing frameworks which consider level of autonomy and risks of AI use. However, these frameworks are coarse and can group very different tools into the same bins with similar validation and governance requirements which may not be appropriate for different AI functions. Application-directed regulation also fits into more nuanced proposed approaches, such as deconstructing clinical work into component activities with distinct standards for automation [15].
For clinicians, the unanswered questions around how new AI tools can be developed, validated and implemented is a call to action. To move beyond image categorisation in glaucoma, clinicians are best placed to identify and describe the areas in which AI tools can improve the care they provide. There are opportunities before, during and following glaucoma consultations for AI tools, but it is up to us to seize them!
References
1. Teo ZL, Thirunavukarasu AJ, Elangovan K, et al. Generative artificial intelligence in medicine. Nat Med 2025;31(10):3270–82.
2. Ling XC, Chen HS, Yeh PH, et al. Deep learning in glaucoma detection and progression prediction: a systematic review and meta-analysis. Biomedicines 2025;13(2):420.
3. Meliante LA, Stuart KV, Luben RN, et al. Current burden and future projections of glaucoma in the United Kingdom. Br J Ophthalmol 2026;110(8):850–6.
4. www.rcophth.ac.uk/wp-content/
uploads/2024/05/Summary-of-RCOphth
-2024-clinical-leads-survey.pdf
5. www.eyenews.uk.com/news/post/introducing
-the-glaucoma-field-defect-classifier
6. Thirunavukarasu AJ, Li S, Qin P, et al. Clinical artificial intelligence applications of vision-language foundation models. PLOS Digital Health 2026;5(6):e0001453.
7. Lima-Cabrita A, Wehbi Z, Pargana J, et al. Artificial intelligence-based glaucoma screening in primary care: a cross-sectional study and economic viability analysis of an ongoing trial. Lancet Primary Care 2026;2(3):100107.
8. Kolovos A, Aung T, Khawaja AP, et al. Polygenic risk scores for glaucoma: Impact on diagnosis and disease course. Prog Retin Eye Res 2026;112:101469.
9. Ravindranath R, Naor J, Wang SY. Artificial intelligence models to identify patients at high risk for glaucoma using self-reported health data in a United States national cohort. Ophthalmol Sci 2024;5(3):100685.
10. Li Z, He Y, Keel S, Meng W, et al. Efficacy of a deep learning system for detecting glaucomatous optic neuropathy based on color fundus photographs. Ophthalmology 2018;125(8):1199–206.
11. Rocha, H. et al. Performance of foundation models vs physicians in textual and multimodal ophthalmological questions. JAMA Ophthalmol 2026;144(1):5–13.
12. Huo B, Boyle A, Marfo N, et al. Large language models for chatbot health advice studies: a systematic review. JAMA Netw Open 2025;8(2):e2457879.
13. CHART Collaborative. Reporting guidelines for chatbot health advice studies: explanation and elaboration for the chatbot assessment reporting tool (CHART). BMJ 2025;390:e083305.
14. Thirunavukarasu A J. How can the clinical aptitude of AI assistants be assayed? J Med Internet Res 2023;25:e51603.
15. Lim E, Thirunavukarasu AJ, He YV, et al. Building a code of conduct for AI-driven clinical consultations. Nat Med 2026;32(2):400–3.
[All links last accessed August 2026]
Declaration of competing interests: None declared.

