Using summary data to detect and quantify ascertainment in biobanks
Ascertainment bias, caused by non-random participation in genetic studies, can distort associations between genetic variants and outcomes. Existing detection methods often require individual-level data, limiting their use on existing summary statistics. The proposed method estimates a parameter, theta, which reflects deviations in the mean polygenic score (PGS) of an ascertained sample relative to the general population. This allows researchers to quantify bias even when only summary data are available, facilitating broader application across biobanks and large-scale consortia. The method could help identify and correct for bias in studies using biobank data, such as UK Biobank, leading to more robust genetic associations. Future work will likely focus on validating the method across diverse ancestries and integrating it into standard GWAS pipelines.