The National Institutes of Health unveiled the world’s most extensive human genome database, a resource designed to accelerate personalized medicine.
Researchers are already creating clinical tools from the NIH’s All of Us program, with polygenic risk scores emerging as a promising method to forecast an individual’s chance of developing complex diseases such as cardiovascular disease and breast cancer—even from early life.
However, these scores are largely trained on DNA from European‑ancestry individuals, leading to poor accuracy for people of other ancestries; for many conditions the predictions for patients of color are barely better than random chance.
Scientists are urgently working to reduce these biases by improving statistical models and enrolling more underrepresented groups, concerned that the gap could widen existing health inequities and hinder the technology’s promise to lower chronic disease rates.
“It really warrants attention,” said Eimear Kenny, the director of the Institute for Genomic Health at Mount Sinai and a principal investigator in a national consortium that tests the genetic scores.
Polygenic risk scores are calculated by aggregating hundreds or thousands of small genetic variations, each contributing a tiny influence on disease risk. When combined, they produce a cumulative score that ranks a person’s predisposition to common diseases.
For individuals at the extreme high end of the risk distribution, a score can provide a medical advantage, giving clinicians decades to intervene—patients with very high cardiovascular risk scores, for example, may begin taking cholesterol‑lowering statins sooner, potentially preventing heart attacks and strokes.
The reliability of these scores depends on the diversity of the data used to train them. The UK Biobank, long hailed as a premier genomic resource, is roughly 94 % white, while the U.S. Veterans Affairs Million Veteran Program is about three‑quarters white.
The NIH’s All of Us program, intended to broaden representation, still comprises roughly half European ancestry participants, according to Broad Institute geneticist Alicia Martin. Moreover, one of its primary funding streams—the 21st Century Cures Act—is slated to expire at the end of the fiscal year, after a 72 % budget reduction since 2023.
“You can’t just snap your fingers and overnight have a biobank from Africa, a biobank from India, a biobank from China that are all as large and open and comprehensive as the UK Biobank,” Dr. Martin said. “We can develop all the fancy statistical methods we want, but still not be able to overcome without more diverse representation.”
Other clinical tools—such as integrated artificial‑intelligence models that combine genomics with biomedical data and tests for rare mutations—also suffer when the underlying genetic databases lack diversity. Doctors sequencing a genome may encounter variants in patients of color that have not been studied sufficiently, limiting interpretation.
Oversights in genetic research can cause scientists to miss breakthroughs that would help everyone. A notable case is the PCSK9 gene, where a rare mutation found primarily in a subset of Black individuals was shown to lower coronary heart disease risk by 88 %. This discovery, which spurred drug development efforts, was made possible only because a Dallas study deliberately recruited a participant pool that was more than half Black.
Demographic history underscores the importance of diverse genetic data. Communities in Africa have accumulated great genetic diversity over hundreds of thousands of years, whereas present‑day Europeans descend from a small group that expanded out of Africa about 50,000 years ago, giving them far less variation.
“If you’re going to pick a population to study for the sake of everybody’s good, the Europeans are the worst you could choose,” said Jay Kaufman, an epidemiologist at McGill University who has written about race and genomics.
The most obvious solution is to recruit more people of color, something the All of Us program says it has prioritized. The initiative works with local partner organizations like churches and community health centers to host discussion sessions and enroll rural communities through a mobile clinic. It also returns individualized health findings and genetic risks to participants to build reciprocity.
Still, recruitment is challenging in the United States, which has a legacy of mistreatment of people of color, from the Tuskegee Syphilis Study to discriminatory sickle‑cell screening and the story of Henrietta Lacks, a Black patient whose cancer cells were used for research without her knowledge or consent.
“Why would you contribute if you’ve been wronged in that way in the past?” Dr. Martin said. “There is some earned mistrust that needs to be addressed—to have that level of trust to be willing to say, ‘Take my data, monitor it for decades, do whatever you want with it.’”
In the meantime, researchers designing polygenic risk scores have found technical ways to diversify data. Modeling tools can now pull weighted information from a variety of sources simultaneously, rather than relying solely on the largest European cohorts and forcing algorithms onto other groups. Mount Sinai Health System in New York and UCLA in Los Angeles, for example, have both established more diverse biobanks than the UK Biobank, each containing tens of thousands of genomes.
Scientists also draw from rapidly growing biobanks in China, Japan, South Korea and Taiwan, which improve the accuracy of polygenic risk scores for East Asian groups. Localized initiatives are underway in Peru, Mexico, Qatar and elsewhere. And a continentwide program called H3Africa is collecting genetic data in African populations and supporting researchers there in studying how genes and environments cause diseases.
Academic networks have tapped sources beyond traditional biobanks, integrating data from large, disease‑specific cohort projects, including the Atherosclerosis Risk in Communities Study (ARIC), which intentionally enrolled thousands of Black participants from North Carolina and Mississippi.
For breast cancer, a project called Confluence aggregates data from over 300 individual studies across 62 countries, providing information from 400,000 breast‑cancer cases and more than 1.5 million controls. The influx of data from minority groups is being used to sharpen breast‑cancer polygenic risk scores for everyone.
Private genomics companies and start‑ups seeking to develop and sell polygenic risk scores typically lacked access to those vast networks. Until recently, they relied almost exclusively on data from the UK Biobank. The All of Us program recently expanded its data‑use policies, allowing companies to draw on its more diverse dataset.
Will these combined efforts be enough? Opinions differ. Some experts argue that even an ethnically representative pool remains imperfect, because socioeconomic status, health‑care access, age or sex can still erode a score’s accuracy. Others contend that the statistical issue was overblown to begin with.
Dr. Kenny, who has been helping to test existing polygenic risk scores for 11 common conditions in racially and ethnically diverse patients, says the bias is too complex to resolve quickly.
“I don’t think these things are perfectly portable yet—and maybe never will be perfectly portable,” she said. “But there are things that are starting to narrow that gap.”


