A preprint describing the reference panel behind the All of Us + AnVIL Imputation Service is now available on medRxiv: A 515,579-Genome Reference Panel Improves Rare-Variant Imputation Across Multiple Underrepresented Populations. The paper describes how the panel was built, how it performs compared to the TOPMed panel, and how it’s deployed through the array imputation pipeline of the All of Us + AnVIL Imputation Service, available at https://allofus-anvil-imputation.broadinstitute.org/.
The largest and most diverse panel in the world
The All of Us + AnVIL reference panel combines 414,830 genomes from the All of Us Research Program with 100,749 genomes from NHGRI’s AnVIL Center for Common Disease Genomics, for a total of 515,579 jointly phased genomes–the largest imputation reference panel produced to date–surpassing the previous record holder, TOPMed R3, which was built from 133,597 participants.
The panel includes 261,163 participants most genetically similar to non-European reference populations. That’s a large enough group that, set apart from the rest of the panel entirely, it would still be the world’s second-largest publicly usable imputation reference panel, behind only the full All of Us + AnVIL panel.
In total, the panel spans 665,398,839 high-quality autosomal sites (~990 million variants), nearly 50% more than TOPMed R3. The complete breakdown of genetically inferred ancestry of the participants in the panel is as follows:
Genetically Inferred Ancestry Group of Participants |
# in reference panel (% of reference panel) |
European (EUR) |
254,416 (49%) |
African (AFR) |
101,982 (20%) |
Americas (previously referred to as “Admixed American,” AMR) |
90,553 (18%) |
East Asian (EAS) |
13,226 (3%) |
South Asian (SAS) |
9,710 (2%) |
Middle Eastern (MID) |
1,065 (0.2%) |
Remaining participants (REM) |
44,627 (9%) |