Heart Failure Risk Scores Across Populations
Heart-failure biomarker scores need validation along two separate axes—within-person behavior over time and performance in populations unlike those used to build them—before their practical meaning is clear.
Heart failure includes biologically and clinically diverse populations, so a risk score can look useful in one dataset yet fail when the population, sampling time, or measurement context changes. Two recent studies examined different parts of that problem. One tracked a metabolic-vulnerability score across a year in a small heart-failure-with-preserved-ejection-fraction cohort. The other compared three protein-based mortality scores, originally created in a community cohort, clinical trials, and a registry, inside one community heart-failure cohort. Together, they show why both temporal behavior and transportability matter, while stopping short of demonstrating that either approach improves decisions or outcomes.
Evidence
The temporal study evaluated the Metabolic Vulnerability Index (MVX), a multimarker score derived from inflammation and malnutrition measures. Investigators used nuclear magnetic resonance profiling on paired baseline and 12-month plasma samples from 46 participants with heart failure with preserved ejection fraction in the TOPCAT trial. Their mean age was 72.5 years; half were women and 89% were White. The study compared the cohort's average MVX and component values at the two time points.
Mean MVX did not change significantly over the year: the mean difference was 1.52 points, with a 95% confidence interval from -0.97 to 4.01 and a p value of 0.225. The authors described the score as appearing stable in this cohort. That conclusion is narrow. A nonsignificant change in the group mean does not show that every participant's score was stable, establish measurement repeatability, or demonstrate that serial change predicts hospitalization or death. The ingested source for this study is abstract-only, so interpretation is limited to the population, methods, result, and conclusion reported there.
The proteomics study addressed a different form of robustness. Researchers recalculated three previously published SomaScan protein risk scores in 1,351 people with heart failure from a southeastern Minnesota community cohort. The original scores had been derived through different settings and methods: a community cohort, clinical trials, and a registry. When applied to the common cohort, the scores were moderately correlated with one another, with Pearson correlations from 0.59 to 0.76.
Each standardized score was associated with all-cause mortality. For each one-standard-deviation increase, the crude hazard ratios were 2.70 for the community-derived score, 1.76 for the trial-derived score, and 1.70 for the registry-derived score. After adjustment for the MAGGIC clinical score and NT-proBNP, the corresponding hazard ratios were 2.40, 1.40, and 1.46. Associations were reported across preserved- and reduced-ejection-fraction subgroups. The three scores also improved time-dependent discrimination beyond the clinical model in this cohort, although the paper did not establish that using the scores prospectively changes care or outcomes.
The protein lists were not interchangeable. Seven proteins appeared in at least two scores, and renin was the only protein included in all three. That overlap is a reproducibility signal worth testing, not proof that renin or any other included protein causes the observed mortality differences. Most selected proteins differed across scores even though overall prognostic performance persisted.
Analysis — Two axes of biomarker reliability
The cross-study connection is methodological rather than a direct comparison. This is analysis, not an established clinical conclusion. The MVX study asks whether a score's average level changes over time in one small trial-derived group; the proteomics study asks whether differently derived scores preserve prognostic associations when brought into one common cohort. Those are complementary axes of reliability. A stable cohort mean can conceal substantial person-to-person movement, while cross-design performance in a single validation cohort can conceal geographic or demographic fragility. Read together, the studies support an emerging research framework in which a heart-failure biomarker score would be tested serially within individuals and independently across health systems, demographic groups, treatment eras, and ejection-fraction categories. Neither paper completed that framework. The apparent convergence is therefore not that one score is ready for use, but that biomarker development must distinguish biological signal from sampling time, cohort selection, and model-construction choices.
Limitations
The MVX analysis included only 46 people, all with preserved ejection fraction, and the cohort was predominantly White. Its primary finding was a comparison of paired population means; the abstract does not report individual trajectories, repeatability statistics, associations between change and outcomes, or power to detect smaller changes. Abstract-only ingestion also prevents independent examination of all component-level and sensitivity analyses.
The proteomics analysis was observational and tested all three scores in the same community cohort, not across several fully independent contemporary cohorts. The community-derived score had been developed using 855 of the 1,351 participants later included in the main evaluation, creating potential optimism; a validation-subset sensitivity analysis reduced but did not eliminate that concern. Source populations were largely older, male, and White, two SomaScan platform versions were used, and the cohorts predated newer heart-failure treatment eras. Hazard ratios and discrimination metrics describe association and prediction, not causal effects or benefit from acting on a score. Larger prospective studies with repeated sampling, prespecified thresholds, independent and more diverse populations, and outcome-focused validation would be needed to clarify practical value.