DiseaseSignal
Proteins & Proteomics

Urine Protein Panels Need Disease Comparators

2026-08-01 · 2 sources · 4 citations · 869 words

Urine proteomics can reveal compact disease-associated protein panels, yet credible diagnostic validation requires realistic clinical comparators, locked models, and independent cohorts.

Evidence

Two primary studies used urine proteomics to search for disease-associated protein panels at different scales. A newly indexed study profiled nine cancer, neurological, and metabolic disease groups against healthy controls. An earlier full-text study focused on endometrial cancer and histologically benign controls. Both reported strong separation inside exploratory case-control datasets, but neither established a clinical diagnostic test.

The July 30 study analyzed 264 urine samples from the Ukraine Association of Biobank. It included 22 participants in each of nine disease groups: kidney, bladder, melanoma, prostate, ovarian, endometrial, and cervical cancers, plus multiple sclerosis and metabolic dysfunction-associated steatohepatitis. Sixty-six healthy participants formed the control group. Laboratory personnel were blinded to disease status, and the investigators measured proteins with the Olink Explore 3072 platform.

The study found three broad expression patterns. Melanoma and endometrial cancer produced relatively few strong, symmetric changes. Cervical, ovarian, and prostate cancers and multiple sclerosis showed more asymmetric patterns with many proteins increased. Kidney and bladder cancers and steatohepatitis showed broad increases. Expected sex-associated differences in proteins including KLK3, MSMB, KLK8, and KLK13 served as an internal check that the assay detected known biology.

For each disease-versus-healthy comparison, multiprotein models outperformed single proteins and generally stopped gaining performance after five to seven proteins. The abstract reports maximum areas under the receiver-operating-characteristic curve of at least 0.95 for seven of nine diseases; multiple sclerosis had the lowest reported maximum at 0.88. It also notes that C9orf40 and pancreatic polypeptide, or PPY, were important across diseases. The authors explicitly described these accuracy estimates as upper bounds requiring prospective validation.

The licensed full-text study examined endometrial cancer more narrowly. Researchers recruited 20 women with histopathologically confirmed endometrial cancer and 20 women with histologically benign tissue at one Saudi center. After an overnight fast, participants provided first-morning midstream urine. The team used label-free liquid chromatography-tandem mass spectrometry rather than an affinity-based panel, quantified 2,662 non-redundant proteins, and applied fold-change and nominal significance thresholds to identify 193 altered proteins: 117 higher and 76 lower in the cancer group.

Internal multivariate analyses separated the two groups. Five-fold cross-validation and 100 permutation cycles were used to examine the supervised model, but all model development remained within this small dataset. Three individual proteins received particular attention. Mitochondrial glutamate dehydrogenase 1 and iduronate 2-sulfatase were lower in the cancer group, with reported within-cohort areas under the curve of 0.945 and 0.965. Histidine-rich glycoprotein was higher, with an area under the curve of 0.875. No independent cohort or orthogonal protein assay tested those candidates.

The endometrial study's controls had benign histology, which is more informative than a healthy-only contrast, but the full text states that it did not include a dedicated group with benign gynecological conditions that could resemble endometrial cancer in practice. The broader study likewise compared each disease with healthy controls rather than demonstrating that one compact panel could distinguish among the nine diseases. Thus, both studies measured case-control separation, while leaving differential diagnosis unresolved.

Analysis — Separation Is Not Specificity

The cross-study inference is that analytical breadth and diagnostic specificity are separate achievements. This is analysis, not a result directly tested across the two datasets. The Olink study shows that a single urine-protein platform can find compact patterns across many disease categories, and the repeated importance of some proteins warns that strong classification against healthy controls may partly reflect signals shared across illness. The LC-MS/MS study independently shows that a focused endometrial-cancer comparison can produce high internal discrimination and a different set of candidates, yet even histologically benign controls do not reproduce the full clinical decision population. Platform, sample processing, population, and model-building choices also differed, so the studies do not replicate a panel. Together they support a staged validation logic: first establish measurable protein differences, then lock the assay and model, and finally test whether the panel separates the target disease from realistic alternatives in an untouched cohort. Performance may fall at each step; measuring that fall is part of validation rather than evidence that discovery failed.

Limitations

The July 30 paper was available to ingestion only as a PubMed abstract. It supports the reported sample structure, platform, pattern categories, and headline accuracy range, but not detailed assessment of missing data, feature selection, model tuning, confidence intervals, or subgroup performance. Each disease group contained only 22 participants, all samples came from one biobank, and the same healthy-control pool underpinned the reported disease comparisons. The abstract does not identify the endometrial-cancer panel or show external validation.

The full-text endometrial study was also small, single-center, and case-control, with 20 participants per group. It mixed cancer stages and grades, lacked an independent validation cohort, and did not confirm candidates with a separate targeted assay. Its internal cross-validation and permutation testing cannot substitute for testing a locked model in newly recruited participants. Neither study shows that urinary protein differences originate from the tumor, predict future disease, improve outcomes, or distinguish the target condition from all relevant alternatives. Their numerical results cannot be pooled because the populations, platforms, preprocessing, and endpoints differed.