Disillusioned by AI in Medicine? There’s a Way Out, but It Might Be Messy

Author: Young May Cha, MD

IARS and SOCCA 2026 Annual Meeting coverage

Artificial intelligence entered medicine with promises of personalized predictions and improved clinical decisions. Despite the large number of published models, however, relatively few have meaningfully changed patient care. At the 2026 IARS and SOCCA Annual Meeting, experts discussed why medical AI remains stuck between promising research and successful clinical implementation, particularly in pediatric perioperative medicine.

Theodora Wingert, MD, described the “messy middle”—the difficult period after a predictive model is developed but before it becomes a reliable clinical tool. Publishing a high-performing model is only the beginning. It must then be validated using real-time data, incorporated into clinical workflows, and continuously evaluated after implementation.

Pediatric models are especially difficult to develop because children represent a smaller and more heterogeneous population than adults. Age-dependent physiology, rapid developmental changes, inconsistent measurements, and relatively uncommon outcomes make validation challenging.

A model may also be statistically accurate but clinically useless. Some systems provide warnings only after clinicians already recognize that the patient is deteriorating. Useful models must generate information early enough to change management and present it in a way that fits naturally into clinical workflows.

Successful implementation requires collaboration among:

  • Clinicians
  • Data scientists
  • Informaticists
  • Statisticians
  • Nurses
  • Hospital administrators
  • Patients and families

Hannah Lonsdale, MBChB, compared AI development with the process used to introduce a new medication. A model should progress through structured phases that include development, external validation, calibration, clinical-impact testing, implementation, and continued surveillance.

Every model should begin with an important clinical question and an adequately large, detailed dataset. Many published systems have high false-positive rates, which can produce unnecessary testing, alarm fatigue, and loss of clinician confidence.

Large multicenter datasets may improve predictive performance, particularly when the outcome being predicted is uncommon. However, artificial intelligence is not always the best solution. Conventional statistical methods may be more transparent, practical, and clinically useful.

Dr. Lonsdale described the NEO-READY model, which uses logistic regression to predict whether a newborn is ready for discharge from the neonatal intensive care unit. This demonstrates that traditional biostatistics can sometimes answer a clinical question more effectively than a complex machine-learning system.

Models must also be monitored for temporal drift. Changes in patient populations, documentation, treatment protocols, and clinical behavior can gradually make a previously accurate model unreliable. Ironically, a successful model may alter clinical practice so effectively that the new data make the model appear to perform worse.

Tori Sutherland, MD, MPH, presented an example involving postoperative opioid prescribing among adolescents. Analysis of national claims data found that teenagers frequently received more than twice the recommended maximum opioid quantity after common procedures, with prescribing levels substantially higher than those reported in other countries.

Prescription records do not prove that patients consumed the medication. Researchers therefore used a new opioid prescription filled three to six months after surgery as a surrogate marker for new persistent opioid use.

Factors associated with increased risk included:

  • Filling an opioid prescription before surgery
  • Age older than 15 years
  • Chronic headache or abdominal pain
  • Depression or anxiety
  • Previous substance use
  • Other mental health diagnoses

Unexpectedly, persistent use was associated with several procedures generally considered less painful, including endoscopy, incision and drainage, and some laparoscopic operations. These findings prompted a review of institutional prescribing practices and increased education for trainees.

Key Takeaways

Medical AI has not failed, but most models have not completed the difficult transition from research publication to routine clinical care.

A high-performance score does not guarantee clinical usefulness. Predictions must arrive early enough to alter treatment and must be presented without creating excessive false alarms or disrupting workflow.

Artificial intelligence should not be used merely because it is fashionable. Traditional statistical methods may be more appropriate when they provide accurate, understandable, and actionable results.

Models require external validation across different institutions and patient populations. Pediatric systems demand particular caution because of developmental variation, smaller datasets, and uncommon outcomes.

Implementation is not the final step. Predictive models must be regularly recalibrated and monitored for changes in clinical practice, patient populations, and data quality.

The path out of AI disillusionment is likely to be slow and complicated. Progress will depend less on developing additional models and more on asking important clinical questions, conducting rigorous validation, integrating tools into actual care, and proving that they improve meaningful patient outcomes.

Thank you to IARS and SOCCA for allowing us to summarize this important coverage from the 2026 Annual Meeting.

Leave a Reply

Your email address will not be published. Required fields are marked *