
Nearly 500,000 participants in the UK Biobank have had their whole-genome sequencing (WGS) data added, making it one of the biggest continuing population-based studies. About 1.5 billion variants—including single nucleotide polymorphisms (SNPs), insertion-deletion (indel) variants, and structural variants (SVs)—have been found in participants of the population-based study. Many of these variants are linked to various disease features and traits and could enable a deeper understanding of disease mechanisms, including those that affect disease risk through non-coding mechanisms.
New dimension of data
The contribution of rare non-coding variation to human diseases and other complex traits has been well-documented. However, research into this area is still in its infancy due to the enormous amounts of data needed from various modalities. Before this work, biological samples and extensive demographic and health-related data had been gathered from 490,640 participants in the United Kingdom as part of the UK Biobank. Subsequent data collection and generation efforts have greatly expanded the dataset’s depth, encompassing fields such as multimodal brain imaging, proteomics, metabolomics, and many more. Previous whole-exome sequencing (WES) on samples from the UK Biobank allowed for characterization of the two-to-three percent exonic genome, but it does not detect SVs and misses almost all non-coding variation.
WGS overcomes the technological constraints of previous genotyping methods by providing a comprehensive and objective picture of the human genome, paving the way for the identification of genetic variation. As part of the most recent data collection wave for the UK Biobank, whole genomes of 490,640 participants were sequenced for an average total of 32.5x base pairs (with a minimum of 23.5x base pairs per individual) using Illumina NovaSeq 6000 sequencing machines. Additionally, 1,175 samples were sequenced twice to ensure quality. Compared to the imputed array, the WGS increased the observed human variation by 18.8 times, while the WES increased it by more than 40 times. Several common variant associations uncovered by WGS that were previously missed by the imputed array data were demonstrated in the study, including those related to hypothyroidism risk and ‘other cataract’ (i.e., ocular disease).
The study also surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights by coupling this dataset with rich phenotypic data. Even though the majority of the associations with disease traits were found in people of European ancestry (93.5% of the total were of non-Finnish European ancestry, while the other 31,785 were from other continental populations), people of African and Asian ancestry also showed strong or new signals.
The non-coding variant problem
We can learn a lot more about how uncommon non-coding variation affects health and disease with this massive, highly phenotyped WGS dataset. Discovering patient populations with specific underlying genetic drivers of disease, validating targets, assessing safety concerns, repositioning opportunities, and answering other questions related to drug discovery and development are all possible with these data. The capacity to prioritize uncommon non-coding variants with a big impact on disease risk will be enhanced thanks to this data’s unique advantage: a better knowledge of the selective constraints operating on disruption outside the coding genome.
Along with the publication of the nearly 500,000 whole genomes, another article was published in Nature Genetics co-authored by Diogo M. Ribeiro, PhD, and Robin J. Hofmeister, PhD, of the University of Lausanne; Simone Rubinacci, PhD, of the University of Helsinki and the Broad Institute; and Olivier Delaneau, PhD, of the Regeneron Genetics Center—analyzing WGS data and 42 blood cell count and biomarker measurements for 166,740 UK Biobank samples. Hundreds of gene-trait associations involving noncoding variants were identified in this second article. The approach identified 911 noncoding gene–trait associations across 42 traits before conditioning on GWAS variants—more than the 723 associations found in coding sequences (CDSs) using the same approach.
While noncoding regions yielded more associations, interpretation proved more complex because coding variants typically affect one gene, but regulatory variants (for example, in enhancers) may influence multiple nearby genes. When associating rare noncoding variants with multiple nearby genes, the results essentially reproduced the results of the nearby genes, suggesting that one must exercise caution in interpreting rare variant associations, which should always be supported by robust conditional analyses. Further research is needed to clarify the role of noncoding rare variants in identifying independent trait associations, including studies on additional traits and larger sample sizes.
So, while the increased scale and richness of the UK Biobank marks a significant milestone in human genetics and provides an improved resource for the exploration of genetic contributions to health and disease, the journey to untangle the complexities of the non-coding genome is still in its early stages.



