
Artificial intelligence (AI) models trained and validated in hospital-based populations to screen for structural heart disease are likely not to be effective in screening within the wider population due to differences in disease prevalence and phenotypes in the real world.
This is according to a report from the PREVUE-VALVE study, a so-called “site-less” clinical trial that recruited subjects aged 65–85 at pharmacies across the USA to screen for undiagnosed aortic, mitral or tricuspid valve disease.
Through the trial, more than 3,000 individuals were signed up in their communities and received an in-home visit where they underwent transthoracic echocardiography (TEE) and 12-lead electrocardiogram (ECG) testing, with the aim of assessing the likely prevalence of moderate or greater valvular heart disease across the population.
Headline findings from the trial were presented at the 2025 Transcatheter Cardiovascular Therapeutics (TCT) conference (25–28 October, San Francisco, USA) by lead investigator David J Cohen (St Francis Hospital and Heart Center, Roslyn, USA), who reported that there may be as many as 4.7 million older adults living with moderate or greater valvular heart disease in the USA, and at least 10.6 million with clinically significant valvular disease, most of whom are unaware of their condition.
In their latest publication from the study programme, released this month in the Journal of the American College of Cardiology (JACC), Timothy Poterucha (Mayo Clinic, Rochester, USA) and colleagues applied the EchoNext (Pathway Labs) AI-ECG model—a previously validated and US Food and Drug Administration (FDA)-cleared AI tool developed by researchers at NewYork-Presbyterian and Columbia University Irving Medical Center (New York, USA)—to the ECG dataset gathered in the trial. Their aim was to evaluate the “transportability” of the AI model from hospital-based to community-dwelling populations and to assess the contribution of disease spectrum and case mix to differences in its performance.
“The EchoNext AI-ECG model was developed in hospital and clinic-based populations undergoing clinically indicated electrocardiography and TTE, reflecting a setting enriched by clinical context,” Poterucha et al note in their JACC paper. “In contrast, PREVUE-VALVE enrolled community-dwelling adults aged 65–85 years with lower disease prevalence and without an indication for clinically driven testing.”
To characterise the differences in disease prevalence and spectrum and their impact on model performance, the researchers compared both disease distribution and model performance in patients aged 65–85 years in the New York-Presbyterian cohort that was
used to train the AI model, alongside other external hospital-based cohorts. They then constructed a propensity-matched sub-cohort of the test population with characteristics similar to PREVUE-VALVE in order assess the extent to which differences in the model performance could be explained by measured differences in patient characteristics.
Compared with hospital-based cohorts, PREVUE-VALVE had lower structural heart disease prevalence (8% vs. 43%), less severe disease, and a shift in phenotype, including more moderate tricuspid regurgitation and less systolic heart failure, Poterucha et al report. Consistent with these differences, model discrimination was lower in PREVUE-VALVE than in the hospital-based cohort, with an area under the curve (AUC) of 71% in the trial compared to 83% for the in-hospital cohort.
Propensity matching attenuated but did not eliminate this difference, the researchers noted, while performance was similar across external hospital-based cohorts, supporting disease spectrum and clinical context as key drivers, they said. The performance of the model was modestly better in PREVUE-VALVE subgroups with higher structural heart disease prevalence and greater disease severity, such as individuals with an abnormal ECG or impaired health status, they also found.
“These findings have important implications for clinical deployment of AI-based diagnostic algorithms such as AI-ECG. In this low-prevalence, community-dwelling population, the positive predictive value of AI-ECG was 14%, meaning that only one in seven positive results reflected true disease,” the study’s authors write. “Even in higher-risk subgroups, such as individuals with abnormal ECGs or impaired health status, positive predictive value increased only modestly to 18% to 26%, which corresponds to a number needed to test of four to seven ECGs to identify a single case of moderate or greater structural heart disease. At scale, this would generate substantial downstream testing, with most evaluations yielding no actionable disease.”
According to the authors, a central question raised by the findings is whether model performance can be improved through retraining or fine-tuning in community- dwelling populations. The AI-ECG model evaluated in the study was developed using more than 700,000 ECGs from clinically enriched populations, enabling it to learn robust patterns associated with more advanced and phenotypically distinct disease.
“Achieving comparable scale in community-based populations with predominantly mild disease may be challenging. Even if such datasets could be assembled, our findings suggest that there may be inherent limitations to model performance in this context,” they note.










