Big data technology lit up on a black background to represent how AI tools like PhyloFrame can help reduce bias in genetic disease prediction
Credit: your_photo / iStock / Getty Images Plus

An equitable machine leaning tool called PhyloFrame developed by researchers at the University of Florida can make research carried out using genetic databases with low diversity more applicable to non-White individuals.

“My dream is to help advance precision medicine through this kind of machine learning method, so people can get diagnosed early and are treated with what works specifically for them and with the fewest side effects,” said lead researcher Kiley Graim, PhD, an assistant professor in the Department of Computer & Information Science & Engineering at the University of Florida, in a press statement. “Getting the right treatment to the right person at the right time is what we’re striving for.”

PhyloFrame is designed to help make precision medicine tools that rely on genetic data, for example, polygenic risk scores for certain medical conditions, more accurate for people from different ancestries. The tool uses disease-related gene expression data, maps showing how the genes interact, and population genomic data to better predict disease outcomes or treatment responses.

The researchers developed the measure of Enhanced Allele Frequency or EAF to identify population-specific enriched variants relative to other human populations in healthy tissue and give a better idea of how they would behave in people with different diseases. “Because EAF is calculated from healthy tissue, this approach can integrate information from under-represented populations that are not present in a smaller disease-specific databases,” the researchers wrote.

To test the model, Graim and colleagues used data from The Cancer Genome Atlas (TCGA), a large cancer study in the United States that includes data from many different types of cancer. The three cancers they chose, breast, thyroid, and uterine, are known to have high ancestral genetic diversity.

For each cancer, the researchers tested PhyloFrame versus a benchmark model on a number of different predictive tests of varying difficulty and varying ancestral bias. “In all three cancers, we show that PhyloFrame mitigates the effects of ancestry imbalances in the training data from the individual disease studies, and improves outcomes for all ancestries,” wrote Graim and colleagues in Nature Communications.

“This holds true when considering the difficulty of prediction tasks, the degree of ancestry bias in the outcome labels, and the ancestry bias in the training data. PhyloFrame signatures are more consistent across models, demonstrating more stability and therefore likely more biological relevance than the comparable benchmark.”

Although this kind of artificial intelligence can help to improve outcomes when diverse disease data is limited, there is no substitute for creating better, more diverse genetic databases in the first place. “Given machine learning modeling relies on large-scale data, it is critical that sequencing efforts rise to this challenge and provide sufficiently diverse data to create equitable precision medicine,” wrote Graim and co-authors.