AI falls short of human experts in assessing severe asthma treatment decisions

Two connected industrial containers with pipes and cables on blue background with stars

A new study comparing five artificial intelligence (AI) platforms to asthma specialists found that AI models frequently disagreed with expert clinicians when recommending biologic treatments for severe asthma, with poor concordance scores ranging from 0.143 to 0.213, suggesting AI is not yet ready to make or support complex asthma treatment decisions.

  • Poor AI performance: Five AI platforms (ChatGPT, Gemini, Copilot, PerplexityAI and DeepSeek) showed weak agreement with specialist recommendations, with concordance scores between 0.143 and 0.213.
  • Major AI errors: Systems produced non-evidence-based recommendations, logical inconsistencies, hallucinations and biased treatment suggestions favoring therapies with larger research footprints.
  • Human advantage: Multidisciplinary teams incorporate perspectives from pulmonologists, nurses and pharmacists to evaluate biomarkers, comorbidities, patient preferences and psychosocial factors that AI cannot quantify.

Recent research raises questions about the readiness of artificial intelligence (AI) to assist with complex medical treatment decisions, finding that leading AI models frequently disagreed with specialist clinicians when recommending biologic therapies for patients with severe uncontrolled asthma. The paper, “Comparison of Artificial Intelligence and Multidisciplinary Team Assessment for Evaluating Biologic Decisions in Severe Uncontrolled Asthma,” was published in The Journal of Allergy and Clinical Immunology.

Scottish researchers compared recommendations from five popular AI platforms (ChatGPT, Gemini, Copilot, PerplexityAI and DeepSeek) against decisions made by a multidisciplinary team (MDT) of asthma specialists. The findings, researchers noted, suggest that AI is not yet ready to serve as either a standalone or supporting decision-maker in this area of respiratory care.

The study analyzed 114 biologic-naïve patients with severe, uncontrolled asthma who were reviewed by the Scottish regional health board NHS Tayside severe asthma MDT between January 2024 and August 2025. Investigators compared the biologic treatment selected by specialists with recommendations generated by the AI platforms. 

Results showed consistently poor agreement between AI recommendations and MDT decisions. Statistical measures of concordance ranged from 0.143 to 0.213, indicating none of the AI models closely replicated expert clinical judgment. Agreement among the AI models themselves was also weak to moderate, suggesting the systems often reached different conclusions when evaluating the same patient information, researchers said.

“AI and large language models are evolving rapidly and are often used in private settings and increasingly in professional activities,” Philipp Suter, MD, told news outlet, Healio. Dr. Suter is a research fellow at the Scottish Centre for Respiratory Research. “Many studies have been conducted in specialties such as radiology, oncology or surgery, but studies assessing AI in respiratory medicine are lacking.” 

Investigators said selecting the right biologic therapy for severe asthma has become increasingly complex as new treatments have entered the market. Specialists must evaluate multiple biomarkers, comorbidities, patient preferences, treatment histories and quality-of-life considerations when selecting among therapies such as benralizumab, dupilumab and tezepelumab, they said.

Unlike AI systems, MDTs incorporate perspectives from pulmonologists, specialist nurses, pharmacists and other healthcare professionals, researchers noted, and these discussions often consider factors that are difficult to quantify, including treatment adherence concerns, psychosocial challenges, frailty and patient-specific circumstances.

The study found that AI models struggled particularly with complex clinical scenarios and frequently produced recommendations that conflicted with current evidence or established guidelines. Researchers identified three major categories of errors: non-evidence-based recommendations, logical inconsistencies and nonsensical outputs. 

One of the most concerning findings, they said, involved AI “hallucinations” and misinformation. Researchers reported that some models cited unsupported claims regarding biologic therapies, overlooked newer clinical evidence or recommended treatments inconsistent with guideline-based care. In several cases, AI systems favored specific therapies despite limited supporting evidence, they noted.

Researchers also observed a tendency for AI models to favor treatments with larger published research footprints, potentially introducing bias. For example, they said, models frequently recommended dupilumab while underutilizing tezepelumab, despite evidence suggesting comparable benefits in certain patient populations.

Despite the disappointing results, researchers emphasized that AI could still play a valuable role in healthcare. Rather than making treatment decisions, they said future AI applications may be better suited to supporting clinicians through literature reviews, workflow management, information synthesis and identification of potential treatment options. 

Researchers called for larger, multicenter studies involving additional asthma MDTs to determine whether the findings can be replicated on a broader scale. For now, they said the study reinforces the importance of human expertise in severe asthma care and suggests that experienced clinicians remain essential to delivering personalized treatment decisions.

More in Asthma
Page 1 of 29
Next Page