Polygenic Scores and Clinical Context
Polygenic scores supplied measurable information across two diseases, yet a compact score matched far larger dementia models and family history remained independently informative for coronary risk.
Evidence
Polygenic risk scores combine the estimated effects of many genetic variants into one measure. Two recent studies tested what those scores add in diverse U.S. cohorts, but they asked different questions. One compared six score-building methods for incident dementia. The other assessed whether a coronary heart disease score, rare familial-hypercholesterolemia variants, and family history add information to a standard clinical risk equation. Together, they show why statistical association and practical prediction are not the same result.
The dementia study analyzed 6,338 participants in the Multi-Ethnic Study of Atherosclerosis who self-identified as Black/African American, Chinese, Hispanic/Latino, or White. Participants were free of dementia at enrollment, had a mean baseline age of about 62, and were followed for a median 16.8 years. Hospitalization and death records identified 560 incident all-cause dementia events, or 8.8% of the sample. The researchers compared three clumping-and-thresholding scores with three Bayesian scores. All scores excluded the APOE region so the analysis could measure polygenic information beyond this major genetic risk locus.
Model size varied enormously. The strictest clumping-and-thresholding score contained 15 variants, while Bayesian scores contained more than 850,000 variants and the largest included 968,595. Yet the compact 15-variant score had an incident-dementia hazard ratio of 1.18 with a 95% confidence interval of 1.08 to 1.28, similar to the 1.17 estimate for the largest cross-ancestry Bayesian score. Adding more variants did not improve discrimination: the highest reported concordance index for a score alone was 0.54, and its area under the curve was 0.55.
The dementia score also added little to established information. A baseline age-specific model containing sex and APOE ε4 status had a concordance index of 0.58. Adding either the strict 15-variant score or the cross-ancestry Bayesian score increased concordance by 0.0069, although both improved statistical model fit. Performance varied across ancestry strata. The 15-variant score was associated with dementia in the lowest and highest non-Finnish-European-like ancestry groups, each with a hazard ratio of 1.27, but no tested score was associated with incident dementia in the intermediate group. No score reached statistical significance in the Chinese subgroup, which had only 40 events.
The coronary heart disease study used two separate datasets: 19,348 participants in eMERGE phase IV and 239,645 in All of Us. Coronary disease covered myocardial infarction, unstable angina, or coronary revascularization. Researchers modeled prevalent disease in eMERGE and incident disease in All of Us, then tested a coronary polygenic score, pathogenic or likely pathogenic variants in familial-hypercholesterolemia genes, and reported family history alongside the pooled cohort equations.
The coronary polygenic score and family history were independently and additively associated with disease in both cohorts, with consistent effects across self-identified White, Black, and Latino groups. In eMERGE, adding both factors to the pooled cohort equations raised the c-statistic from 0.719 to 0.753. At the 7.5% 10-year-risk threshold, 18.8% of participants were reclassified, corresponding to about four additional true-positive coronary disease identifications per 1,000 people screened. The abstract also reports net-benefit gains between 7.5% and 10% thresholds across all three groups.
Analysis — Context Determines Added Value
The cross-study pattern is an analysis, not a clinical rule: a polygenic score matters only relative to the information already available and the decision it is meant to improve. In the dementia study, scores were statistically associated with future events, but discrimination remained low and adding a score to sex plus APOE changed concordance only slightly. More computational complexity did not solve that problem; 15 selected variants performed comparably to models with hundreds of thousands. In the coronary study, the score produced a clearer incremental signal when combined with a validated clinical equation and family history, including measurable reclassification and decision-curve gains. Family history did not become redundant after genotyping, which suggests that it captures inherited, environmental, and shared-behavior information not fully represented by the score. The emerging direction is therefore not “more variants equals better prediction.” It is careful evaluation of whether a score adds calibrated, equitable information beyond existing clinical variables in a defined population and outcome.
Limitations
These studies cannot be compared as if they were one validation experiment. They examined different diseases, endpoints, baseline models, score construction methods, and performance measures. The coronary evidence available in the ingested pack is abstract-only, so subgroup effect sizes, familial-hypercholesterolemia variant counts, calibration details, missing-data procedures, and the full decision-curve analysis cannot be assessed here. Its eMERGE analysis was cross-sectional for prevalent disease, while All of Us supplied the incident analysis.
The dementia study relied primarily on diagnosis codes from hospitalizations and death records, which may miss cases and do not isolate Alzheimer disease from other dementias. Alzheimer-focused genetic scores were therefore tested against an all-cause dementia outcome. The Chinese subgroup and ancestry-stratified event counts limited precision, and the available training genome-wide association studies underrepresented Hispanic/Latino and Chinese ancestry. Even the best score-only concordance was below 0.60. Neither study shows that a polygenic score determines an individual's future or establishes routine clinical benefit. Prospective evaluations with prespecified actions, calibration checks, broader ancestry representation, and outcome-specific follow-up would be needed to establish whether the reported statistical gains improve care.