From Technical Validation to Real-World Impact: The HFpEF Case

Ashley Schmidt VP | GM AHA AI Assessment Lab

Healthcare rarely produces findings where every stakeholder wins. Patients, providers, payers, and pharmaceutical developers each operate within different incentive structures, optimizing for different outcomes, often at each other's expense. Interventions that genuinely align those interests, rather than trade them off, are unusual enough that when evidence of such alignment emerges, it is worth examining carefully.

An Impact Analysis we recently conducted through the American Heart Association AI Assessment Lab, powered by Dandelion Health, produced exactly that kind of evidence. We applied a rigorous retrospective assessment to an FDA-cleared AI algorithm that analyses routine echocardiograms to identify heart failure with preserved ejection fraction (HFpEF); a condition that affects millions of people and goes undetected (and thus untreated) in approximately two thirds of them under current clinical workflows. Our Impact Analysis quantifies the actual real-world clinical and economic impact of an algorithm once it is deployed over years of time, not just validating an algorithm’s technical feasibility. 

What the Impact Analysis found

The algorithm detected HFpEF a median of 263 days earlier than standard care. That gap is not a marginal efficiency gain, it is the difference between treatment initiated before a patient decompensates, and treatment initiated in response to an acute event.

Earlier detection in this dataset translated into meaningfully higher rates of guideline-directed therapy: mineralocorticoid receptor antagonist use was 1.8 times higher and began 247 days earlier; SGLT2 inhibitor use 1.7 times higher and 283 days earlier; ARNI use 3.3 times higher and 276 days earlier. 

When you can see the treatment uptake change in the data, the downstream economic projections become substantially more credible than projections that assume treatment effects without demonstrating them:

477 fewer deaths

406 fewer hospitalisations

$1.9M in additional net revenue for every 10,000 patients over five years

One finding deserves particular attention: cost-effectiveness was highest in non-white patients, where diagnostic delay tends to be most pronounced. That is not a secondary result. It is central to what rigorous evaluation on real-world populations is supposed to surface.

Why this kind of evidence is different 

Most clinical AI is evaluated on technical benchmarks. These tell you whether an algorithm performs against a test set. They do not tell you whether deploying it changes what happens to patients. The gap between technical validation and demonstrated impact is where most clinical AI currently sits and it is the gap that makes health systems, payers, and procurement committees hesitant.

The assessment ran on Dandelion’s proprietary real-world dataset, which links clinical records, echocardiograms, prescriptions, and long-term outcomes across a large health-system population. When we traced what happened to patients whose HFpEF was identified early versus those whose diagnosis was delayed, following both groups forward through years of clinical and economic outcomes in a longitudinal real-world dataset, the findings pointed in the same direction for every group in the system. 

Patients were treated before they could decompensate. Health systems reduced their exposure to readmission penalties under CMS frameworks. Payers avoided the highest-cost acute utilization events. And for pharmaceutical developers, a population that is currently largely invisible in standard real-world data becomes understandable for planning trial design, recruitment, and launch planning.

What this suggests about clinical AI more broadly

Not all algorithms will produce this kind of convergence. It depends on the condition, the diagnostic gap, and the data available to trace downstream consequences. In this analysis, the signal was consistent and the methodology was rigorous enough to make the findings credible. Whether that convergence holds in other settings, other conditions, other algorithms, other health system contexts, is an open empirical question. But the framework for asking it, and the infrastructure to answer it with the depth the question deserves, is increasingly available. That is what this study, in our view, begins to demonstrate.

This Impact Analysis was accepted to ISPOR 2026, the main meeting for health economics, where it placed in the top one percent of submitted abstracts.

Sample Impact Analysis output clinical outcomes from the Ultromics HFpEF use case

Next
Next

Cardiovascular Disease Is Entering Its Precision Era. Trial Design Needs to Catch Up.