Clinical Research bioRxiv (all subjects)

ARCHER-LD: Rapid Long-Range Linkage Disequilibrium Calculations at Biobank Scale using GPU Acceleration

linkage disequilibriumGPU accelerationbiobank-scale genomicsfine-mapping

LD information from a study's own dataset is optimal for downstream analyses such as statistical fine-mapping, but computing all pairwise LD is computationally expensive: for N variants it requires roughly N^2/2 computations. This has led most studies to use external reference panels such as 1000 Genomes, which is a limitation for biobank-scale whole-genome sequencing datasets containing hundreds of millions of variants.

The authors present ARCHER-LD, a multi-GPU distributed computing approach that calculates R^2 for every variant pair in a dataset. On chromosome 22 of the 30x WGS 1000 Genomes dataset (1.8 million variants), running on eight consumer-level 12GB NVIDIA RTX 2080Ti GPUs took 80 minutes, an approximately 8x speedup. In the Penn Medicine Biobank (57,170 samples), chromosome 1 LD (~1.38 million variants) was computed in 49 minutes versus 22.7 hours for PLINK, an approximately 28x speedup. For full biobank-scale demonstration, they computed the entire LD matrix (>1e16, or 10 quadrillion elements) for the 30x WGS 1000 Genomes dataset (~120 million variants, ~2,500 samples) in under 6 hours using 512 NVIDIA 40GB A100 GPUs on the Department of Energy Argonne Leadership Computing Facility Polaris Supercomputer.

The tool is coded in Python using CuPy and is publicly available. By enabling researchers to compute LD from their full genomic data instead of external panels, it can yield more accurate, population-specific findings, particularly for groups underrepresented in existing LD reference databases.

Read original →

← Back to home