On-demand Webinar: Clinical Blended Genome-Exome Sequencing in Practice

New preprint available: A 515,579-genome reference panel powers the All of Us + AnVIL Imputation Service

Sofia Labrecque

A preprint describing the reference panel behind the All of Us + AnVIL Imputation Service is now available on medRxiv: A 515,579-Genome Reference Panel Improves Rare-Variant Imputation Across Multiple Underrepresented Populations. The paper describes how the panel was built, how it performs compared to the TOPMed panel, and how it’s deployed through the array imputation pipeline of the All of Us + AnVIL Imputation Service, available at https://allofus-anvil-imputation.broadinstitute.org/.
The largest and most diverse panel in the world
The All of Us + AnVIL reference panel combines 414,830 genomes from the All of Us Research Program with 100,749 genomes from NHGRI’s AnVIL Center for Common Disease Genomics, for a total of 515,579 jointly phased genomes–the largest imputation reference panel produced to date–surpassing the previous record holder, TOPMed R3, which was built from 133,597 participants.
The panel includes 261,163 participants most genetically similar to non-European reference populations. That’s a large enough group that, set apart from the rest of the panel entirely, it would still be the world’s second-largest publicly usable imputation reference panel, behind only the full All of Us + AnVIL panel. 
In total, the panel spans 665,398,839 high-quality autosomal sites (~990 million variants), nearly 50% more than TOPMed R3. The complete breakdown of genetically inferred ancestry of the participants in the panel is as follows: 
Genetically Inferred Ancestry Group of Participants
# in reference panel (% of reference panel)
European (EUR)
254,416 (49%)
African (AFR)
101,982 (20%)
Americas (previously referred to as “Admixed American,” AMR)
90,553 (18%)
East Asian (EAS)
13,226 (3%)
South Asian (SAS)
9,710 (2%)
Middle Eastern (MID)
1,065 (0.2%)
Remaining participants (REM)
44,627 (9%)
Note: this information can be found in Fig. 1A and Supplementary Table 1 in the manuscript.
Rare-variant accuracy, validated against whole-genome sequencing
The preprint validates the panel against 42 samples with matched genotyping array and 30x whole-genome sequencing data, evaluated across six genetic ancestry validation groups: two groups most similar to the African reference population but drawn from different source populations (labeled AFR-SoS and AFR-US in the preprint, correlating to African-South-of-the-Sahara and African-American reference populations, respectively), plus groups most similar to the AMR, EAS, EUR, and SAS reference populations. 
At a standard imputation quality filter (inferred R² > 0.3), imputed SNVs reached high accuracy (empirical R² > 0.8) at allele frequencies as low as 0.2% across 5 of 6 ancestry groups tested, extending reliable imputation deep into the rare-variant frequency range. Imputed indels also achieved high accuracy, with (empirical R² > 0.8) at allele frequencies as low as 2% across 5 of 6 ancestry groups tested. 
Compared against TOPMed on the same validation samples, the All of Us + AnVIL panel improved imputation accuracy in 5 of 6 ancestry groups, with gains of more than 10% in empirical R² for rare variants (0.1% minor allele frequency) among AMR, EAS, EUR, and SAS samples.
To evaluate downstream impact, the panel was also checked against 243 trait-associated variants from a UK Biobank whole-genome sequencing study that had been missed by TOPMed imputation. The All of Us + AnVIL panel uniquely recovered 20 of these variants, including 7 rare variants not present in TOPMed.
Secure and scalable infrastructure
The reference panel is de-identified before deployment using RESHAPE, a method that simulates multiple generations of genetic recombination to produce privacy-preserving recombined haplotypes. End users do not have access to All of Us or AnVIL individual-level genotype data (which comprises the reference panel), so they do not need to apply for access to the data to leverage it via the Imputation Service.
The Imputation Service runs on Terra with FedRAMP/NIST 800-53 Rev. 5 Moderate compliance, supporting use with controlled-access data. Users submit jobs through the command-line tool or the graphical interface at https://allofus-anvil-imputation.broadinstitute.org/. In testing, runtime scaled nearly linearly with sample size: submissions of 2,000, 5,000, and 10,000 array samples completed in 22, 42, and 75 hours, respectively.
What’s next
The All of Us + AnVIL Imputation Service now offers array and low-pass whole-genome sequencing (including blended genome-exome) imputation pipelines, all built on this same reference panel. A third pipeline, for imputing structural variants, is in development and is on track for release fall 2026, extending the service beyond SNVs and indels to a broader range of genetic variation.
Read the preprint
The full preprint is available on medRxiv at https://www.medrxiv.org/content/10.64898/2026.08.25.26361247v1. To impute your own array data against the panel, visit https://allofus-anvil-imputation.broadinstitute.org/.

Return to the Blog

Sean Hofherr

Chief of Clinical Strategy and Product Development, Broad Clinical Labs

Sean Hofherr is dual board certified by ABMGG in Clinical Biochemical Genetics and Clinical Molecular Genetics. Sean serves as the Chief of Clinical Strategy and Product Development at Broad Clinical Labs. In this role at BCL, Sean is able to leverage his extensive experience to guide the clinical vision and delivery across the organization. Sean most recently served as the Chief Operating Office at Fabric Genomics, which focuses on the use of AI and Bioinformatics for Clinical Interpretation of whole genome sequencing. Prior to Fabric, Sean was the Chief Scientific Officer and CLIA Director at the commercial reference laboratory, GeneDx.

Sean received his B.S. degree in Microbiology and Cell Sciences from the University of Florida before earning his Ph.D. in Molecular and Human Genetics from Baylor College of Medicine. Sean completed clinical fellowships in Clinical Biochemical Genetics and Clinical Molecular Genetics at the Mayo Clinic.

Danielle Perrin

Chief of Staff, Broad Clinical Labs

As Broad Clinical Labs’ Chief of Staff, Danielle Perrin advises and supports colleagues on the executive leadership team in BCL’s strategic planning and execution. She builds and leads new organizational functions and processes and leads critical projects, as well as driving effective information flow, decision making, and execution throughout the organization. An operations leader with a business, engineering, and biology background and 20+ years of experience in the genomics field, Perrin has a track record of driving operational excellence and building and scaling both physical and business processes. During her career at Broad, which started in 2003 at the tail end of the Human Genome Project, Perrin has led laboratory operations and R&D teams in Broad’s Genomics Platform, as well as fulfilling senior advisory and leadership roles in the Broad Institute’s COO and CFO offices.

Perrin received her B.S. in Biology and M.E. in Biotechnology Engineering from Tufts University and her M.B.A. from the MIT Sloan School of Management.

Tim De Smet

Chief Commercial Officer, Broad Clinical Labs

As Chief Commercial Officer of Broad Clinical Labs, Tim De Smet leads BCL’s business development, alliance management, external project management, and customer support teams. A Broad Institute employee since 2008, De Smet has held leadership roles and managed teams of various sizes in Broad’s Genomics Platform and clinical lab, spanning laboratory operations, finance, and informatics, and has expertise in work design, financial modeling, and high scale laboratory and business operations.

De Smet received his B.S. in Biochemistry and M.B.A. from Northeastern University.

Jim Meldrim

Chief Technology Officer, Broad Clinical Labs

As Chief Technology Officer, Jim Meldrim sets the vision for Broad Clinical Labs’ informatics systems, including the hardware and software used for sample intake and tracking, data production, analysis, and delivery. Having held a variety of laboratory and informatics-focused leadership roles at Broad, spanning R&D and production operations, Meldrim has been a leader and innovator in the generation, management, and analysis of genomic data since 1999, beginning with sequencing data generation for the Human Genome Project.

Meldrim received his B.S. in Biology from Cornell University.

Sheila Dodge

Chief Operating Officer, Broad Clinical Labs

As Chief Operating Officer, Sheila Dodge leads Broad Clinical Labs’ process development and implementation activities, as well as lab operations, financial planning and operations, quality & compliance, and core business processes. A Six Sigma Black Belt with extensive experience in process development and high throughput genomics operations, Dodge is an expert in work design and in collaborating with a range of collaborators, scientists, engineers, and technology partners to rapidly integrate new technologies and operationalize innovations. A member of the Broad Institute since 2001, Dodge is an Institute Scientist and lectures at the MIT Sloan School of Management on operations, dynamic work design, and visual management techniques.

Dodge received her B.A. in biochemistry and molecular biology from Boston University and her master’s degree in biology from Harvard University. She earned her M.B.A. from MIT Sloan School of Management.

Heidi Rehm, Ph.D., FACMG

Chief Medical Officer and Clinical Laboratory Director, Broad Clinical Labs

Heidi Rehm is board-certified by ABMGG in Clinical Molecular Genetics and Genomics and serves as BCL’s Chief Medical Officer and Clinical Laboratory Director. She oversees BCL’s regulatory requirements, leads the clinical team performing genomic interpretation and variant analysis, and guides BCL’s efforts in genomic testing for clinical and research use. She is also an Institute Member of the Broad and co-director of the Medical and Population Genetics Program. Rehm is also the Chief Genomics Officer in the Department of Medicine and Genomic Medicine Unit Director at the Center for Genomic Medicine at Massachusetts General Hospital, working to integrate genomics into medical practice. She is a principal investigator of ClinGen, providing free and publicly accessible resources to support the interpretation of genes and variants. She co-leads both the Broad Center for Mendelian Genomics, focused on discovering novel rare disease genes, and the Matchmaker Exchange, which aids in gene discovery. She is Chair of the Global Alliance for Genomics and Health, a principal investigator of the Broad-LMM-Color All of Us Genome Center, co-leader of the Genome Aggregation Database (gnomAD), and a Board Member and Vice President of Laboratory Genetics for the American College of Medical Genetics and Genomics.

Rehm received her B.A. degree in molecular biology and biochemistry from Middlebury College before earning her M.S. in biomedical science from Harvard Medical School and Ph.D. in genetics from Harvard University. She completed her post-doctoral training with David Corey in neurobiology and a fellowship in clinical molecular genetics at Harvard Medical School.

Niall Lennon, Ph.D.

Chair and Chief Scientific Officer, Broad Clinical Labs

As Chair and Chief Scientific Officer of Broad Clinical Labs, Niall Lennon leads the team and sets the scientific and clinical vision for the organization. Dr. Lennon joined the Broad Institute in 2006 and has since contributed to the development of applications for every major massively parallel sequencing platform across a range of fields. In 2013 Dr. Lennon led the effort to establish a CLIA licensed, CAP-accredited clinical laboratory at the Broad Institute to facilitate return of results to patients and to support clinical trials. More recently, he has led efforts to achieve FDA approval for large-scale genomics projects (NIH’s All of Us Research Program) and for Broad’s own clinical diagnostic for COVID-19 testing operation, which returned 37+ million results to patients. Dr. Lennon is a principal investigator of the eMerge and All of Us projects, an Institute Scientist at Broad, Associate Director of Broad’s Gerstner Center for Cancer Diagnostics, and an adjunct professor of biomedical engineering at Tufts University, where he teaches Molecular Biotechnology.

Dr. Lennon received a Ph.D. in pharmacology from University College Dublin and completed his postdoctoral studies at Harvard Medical School and Massachusetts General Hospital. He holds an executive certificate in management from the MIT Sloan School of Management.