Hundreds of AI tools are entering clinical care, but researchers traced how quickly the evidence thins out when the focus shifts from regulatory authorization to trials that measure what actually happens to patients.

Study: 1,357 AI medical devices cleared, 3 actually tested on patient outcomes. Image Credit: Iryna Pohrebna / Shutterstock

A recent study published in the journal PLOS Digital Health found a significant gap between the number of artificial intelligence (AI)-based medical devices cleared or approved by the United States Food and Drug Administration (US FDA) and the number evaluated for patient outcomes.

Of 1,357 FDA-cleared or approved AI/ML-enabled devices, only three had been evaluated for patient-centered outcomes. Whether these devices can meaningfully improve patients’ lives remains largely unanswered based on publicly available patient-outcome evidence. If effectiveness were confirmed in further high-quality research using standardized workflows and diverse populations, clinicians could use these tools more confidently in routine practice to improve patient care.

AI tools are rapidly entering clinical practice worldwide, with 1,357 AI/ML-enabled devices cleared or approved by the US FDA for applications spanning radiology, cardiovascular medicine, and neurology.

Despite widespread availability, most of these tools have not been rigorously evaluated in routine clinical practice. Although many can reach the market by demonstrating substantial equivalence to existing predicate devices, their actual benefit to patients remains unclear.

Funnel diagram showing attrition of clinical evidence for FDA-cleared AI devices. Each stage displays the count and percentage relative to all cleared devices (1,357). Tapered widths reflect the declining number of devices with registered trials, posted results, peer-reviewed publications, and patient outcomes.

About the study

In the present study, researchers investigated the publicly available clinical validation evidence for more than 1,300 FDA-cleared or approved AI devices. They searched the American College of Radiology (ACR) Data Science Institute (DSI) catalog and the FDA device database for relevant records through 5 December 2025. In addition, they identified peer-reviewed records published in PubMed and clinical trials registered on ClinicalTrials.gov.

In the study, “FDA-cleared” described devices that had gone through the 510(k) pathway, while “FDA-approved” applied to devices authorized under the Premarket Approval (PMA) pathway. The broader term “FDA-authorized” encompassed devices that reached the market through the 510(k), De Novo, and PMA pathways.

Reviewers from engineering or clinical research backgrounds extracted data on study characteristics, developer type, participant demographics, exclusion criteria, and sample size, resolving discrepancies by consensus.

The analysis covered patient-relevant outcomes, including hospitalizations, readmissions, quality of life, functional status, and symptom burden. Major events such as myocardial infarction, stroke, serious adverse events, and death were also considered. The EuroQol Five Dimensions (EQ-5D) and 36-Item Short Form Health Survey (SF-36) scales provided examples of standardized measures of quality of life. Disability scores, activities of daily living (ADLs), and return to work helped assess the potential impact of the tools on participants’ functional status.

Diagnostic accuracy indicators such as sensitivity, specificity, and physiological measurements were considered surrogate endpoints when they had no clear link to patient outcomes.

For evidence mapping, the team searched PubMed using device names, manufacturer details, and outcome terms. They then checked the FDA’s 510(k) summaries for corresponding trial registration numbers and used NCT numbers to connect the records.

Results and discussion

The analysis revealed a striking gap between regulatory clearance and clinical validation. Fewer than 3% (n=34) of 1,357 FDA-cleared devices were linked to registered clinical trials.

Surprisingly, not even one percent (n=12) of the devices had results posted to ClinicalTrials.gov, while 12 had peer-reviewed publications, and only three devices (0.2%) evaluated patient outcomes such as morbidity, readmissions, or mortality. Most available evidence instead relied on retrospective accuracy measures or surrogate endpoints. These datasets may not reflect the diversity of patient populations across the globe.

Most studies were observational (62%), while 68% of the 34 registered trials were conducted exclusively in the United States. The authors suggest this gap may partly reflect commercial and regulatory incentives that favor rapid market entry. On the other hand, prospective, multi-center trials are usually costly, time-consuming, and difficult to conduct without skilled personnel with clinical expertise to integrate electronic medical records into clinical workflows. Since AI technology is evolving rapidly and trials run for months to years, the version under investigation may become outdated before the trial is completed.

In external validation studies cited by the researchers, the Epic Sepsis Model missed 67% of sepsis cases while triggering alerts in 18% of hospitalizations. IBM Watson for Oncology showed concordance with expert recommendations as low as 12% for gastric cancer in China and 33% in Denmark, highlighting risks of insufficient context-specific validation.

To address existing gaps, the researchers propose a staged framework that begins with retrospective validation on diverse, representative datasets. This would be followed by prospective workflow-embedded studies involving at least 500 individuals with pre-specified safety and usability endpoints, and multi-center trials of at least 2,000 participants evaluating patient-centered outcomes and relevant population subgroups.

The authors caution that the analysis was limited to publicly available evidence and may have missed proprietary or unregistered validation studies not linked in the databases examined. Some FDA 510(k) summaries were also unavailable, and the study did not include a non-AI device comparison group, meaning the identified evidence gaps may not be unique to AI devices.

Conclusion

The findings highlight an enormous unfilled gap between regulatory authorization and clinical validation of AI technology. With only 0.2% of FDA-cleared devices evaluated for patient-centered outcomes, publicly available patient-outcome evidence has not kept pace with regulatory authorization.

In addition, while these devices are primarily evaluated in high-income settings such as the US, their use in resource-limited settings, even after regulatory authorization, may not translate into meaningful benefits for underserved populations. 

In future studies, the authors propose staged validation, progressing from preliminary investigations with limited sample sizes to prospective trials across population subgroups. 

Greater transparency and stronger ethical safeguards would allow clinicians, health systems, and patients to assess the benefits and risks of healthcare AI more independently.