Clinical Research medRxiv (all subjects)

Generalizability of proteomic risk prediction across biobanks reveals dependence on phenotype definitions

proteomicsrisk predictionUK Biobankgeneralizability

Advances in high-throughput proteomics now allow health states to be assessed in biobank-scale cohorts. Disease prediction models built on these data have outperformed baseline clinical models across many diseases and can offer insight into disease pathogenesis, but whether they generalize across cohorts had not been established at scale.

The study trained models for 15 diseases in UK Biobank (n=53,026) using Olink proteomics data. The models achieved high disease prediction accuracy (mean AUC=0.74; range 0.56–0.89) and improved on clinical-factor-only models by a mean ΔAUC of 0.03. Performance did not depend on model architecture: simpler models such as L2 performed as well as complex models such as transformers.

Generalizability was assessed in two external cohorts, FinnGen (n=5,865) and All of Us (n=7,405), spanning two Olink platforms. Proteomics-based prevalent (classification) and incident (prediction over the next 5 years) disease models largely generalized across cohorts, although performance varied by disease. Specifically, 10 of 14 prevalent models and 13 of 15 incident models showed no significant performance decrease in any biobank.

After adjusting for demographic, ancestry, and technical covariates, the authors found that differences in phenotyping quality were likely the major drivers of cross-cohort variability. They conclude that proteomic risk models can be powerful and generalizable predictors of disease across multiple cohorts.

Read original →

← Back to home