NIH Expands All of Us Dataset with 747,000 Participants' Genomics and Health Data
NIH expanded its All of Us dataset, linking 535,000 whole genome sequences to 482,000 EHRs from 747,000 participants. The release includes 1.3 billion variants and aims to advance precision medicine, prioritizing underrepresented populations.
The National Institutes of Health (NIH) has released a major expansion of its All of Us Research Program dataset, providing researchers with access to genomic, clinical, behavioral, environmental, and wearable-device data from more than 747,000 participants. The updated resource links nearly 482,000 electronic health records to more than 535,000 whole genome sequences, creating what the NIH described as the largest integrated source of such data in the world.
The new data release, announced on 30 June, was called "a combination of genomic depth and clinical breadth, unmatched by any research programme in the world." It contains more than 1.3 billion genetic variants, allowing researchers to study common and rare genetic differences alongside diagnoses, laboratory findings, medication use, socioeconomic factors, and environmental exposures.
All of Us launched nationally in 2018. Participants may contribute biospecimens, electronic health record information, physical measurements, survey responses, and data from wearable devices. These data are organized within a secure, cloud-based research environment known as the All of Us Researcher Workbench. The program aims to enroll at least 1 million participants and follow their health over time.
A central feature of the program is its emphasis on populations historically underrepresented in biomedical research. More than 86% of participants represent communities that have traditionally been overlooked, including racial and ethnic minority groups, rural populations, individuals with disabilities, and people facing socioeconomic barriers. This contrasts with global genomic databases, where nearly 80% of participants in genome-wide association studies are of European descent; in the NHGRI-EBI GWAS Catalogue, only 2.4% of participants are of African ancestry, yet that 2.4% contributed to 7% of findings regarding the association of genetic variation with outcomes.
The program has also faced gaps in real-world data. Although 98% of participants agree to share their electronic health records, more than 300,000 have no EHR data in the database. In its latest data release, All of Us is attempting to fill those gaps through an innovative use of patient data-sharing networks primarily used to coordinate clinical care, specifically the eHealth Exchange health information network.
Linking genomic findings with clinical and real-world data may help researchers evaluate the combined effects of biology, behavior, environment, and health care access. The resource could support pharmacogenomic research, including gene-drug relationships, medication-associated adverse events, treatment adherence, prescribing patterns, and health disparities. An earlier analysis of 245,388 whole-genome sequences from All of Us identified more than 1 billion genetic variants, including over 275 million that had not previously been reported.