DiseaseSignal
Genetics & Genomics

Rare Disease Genome Reanalysis

2026-07-23 · 2 sources · 4 citations · 816 words

For these cohorts, updated interpretation of existing genome and exome data contributed more confirmed diagnoses than adding a newer sequencing technology, although long-read sequencing may still matter in other variant classes and populations.

Evidence

A negative genome or exome result is not necessarily a permanent result. The sequence data can be revisited as gene-disease links, population-frequency resources, phenotype descriptions, and variant-calling methods change. Two primary studies tested this idea in different groups of previously unsolved families. One focused on neuromuscular disorders and reanalyzed existing next-generation sequencing data. The other added long-read genome sequencing to a small pediatric rare-disease cohort while also reexamining prior short-read results.

The July 20 neuromuscular study reanalyzed data from 101 undiagnosed families using the RD-Connect Genome-Phenome Analysis Platform. The starting data included clinical exomes from 45 families, whole exomes from 31, and whole genomes from 25. Researchers prioritized variants using population frequency, Human Phenotype Ontology terms, computational predictions, and genotype-phenotype matching. This was therefore not simply another pass through unchanged filters; it combined updated bioinformatics with expert interpretation of how each candidate fit the observed disorder.

Reanalysis identified causative variants in 17 of the 101 previously unsolved families, a reported diagnostic yield of 16.83%. Eight cases involved coding variants in known neuromuscular-disease genes. Five involved intronic variants in known genes, supported by computational predictions and phenotype correlation. The study also reported an expanded phenotype associated with PTPN11, a dual diagnosis involving MYH2 and KIF21A, and the same novel ATP2A2 missense variant in two unrelated families. The authors interpreted the ATP2A2 finding as evidence for a new neuromuscular-disease gene, but independent replication and functional work remain important for establishing the breadth of that relationship.

The second study enrolled 20 families with children or young adults who had suspected rare genetic disorders after negative or inconclusive clinical short-read testing. Nineteen families underwent PacBio HiFi long-read genome sequencing; one family with limited DNA received short-read genome sequencing only. Eleven probands also received research short-read genome sequencing. The researchers used phased-assembly and read-based long-read pipelines, a separate short-read pipeline, phenotype-guided prioritization, population-frequency filtering, and review with the medical team.

Likely pathogenic or pathogenic variants produced diagnoses in 2 of 20 families, a 10% yield. Five more families had findings of uncertain diagnostic significance, not confirmed diagnoses. Crucially, every reported variant was also detected independently through research short-read sequencing or reanalysis of earlier clinical short-read data. The full text states that the diagnostic findings were small variants rather than structural variants unique to long-read sequencing.

Long-read data nevertheless exposed the scale of the interpretation problem. Population-frequency filtering reduced structural-variant calls by 66.2% to 88.9%, depending on the caller. Manual review then found many prioritized calls to be technical artifacts, especially in repetitive or otherwise difficult genomic regions. The result was not that long reads failed technically: the pipelines detected known structural variants in positive controls. Rather, in this small affected cohort, the additional variant visibility did not produce a long-read-specific diagnosis.

Analysis

The cross-study pattern is that diagnostic value can change even when a person's underlying DNA sequence does not. This is analysis, not a claim that reanalysis will resolve every unsolved case. In the neuromuscular cohort, updated phenotype matching and variant interpretation recovered coding and intronic findings from existing datasets. In the pediatric cohort, a newer sequencing technology generated extensive structural-variant information, yet the confirmed diagnoses were recoverable from short-read data once those data were reexamined. Together, the studies suggest that the bottleneck can sit downstream of sequencing: candidate filtering, gene-disease knowledge, family segregation, phenotype detail, and expert review determine whether a detectable variant becomes a credible diagnosis. A testable workflow hypothesis is that scheduled reinterpretation and carefully defined escalation to newer sequencing methods may outperform a technology-first approach for some unsolved cohorts. That direction remains unproven across rare disease as a whole, because both studies were selected cohorts and only one directly compared long- and short-read evidence.

Limitations

These studies cannot establish a universal diagnostic strategy. The neuromuscular report covered a larger but disease-specific cohort, and the ingested evidence available here was its PubMed abstract rather than commercially reusable full text, limiting independent inspection of methods and case-level evidence. Its ATP2A2 result came from two unrelated families and needs broader replication and functional validation before the full gene-disease relationship can be considered settled.

The long-read study was small, with 20 families, and its participants had already undergone heterogeneous prior testing. Its absence of a long-read-specific diagnosis does not show that long-read sequencing lacks value in other cohorts, especially where repeat expansions, complex structural variants, methylation, or hard-to-map regions are plausible. Five families had uncertain findings that require functional work, and the high artifact burden shows that detection is not equivalent to pathogenicity. Neither study was randomized, diagnostic yield depends on referral and interpretation practices, and neither result supports predictions about any individual case.