What a Polygenic Score Can and Can’t Tell You From Your 23andMe File
A polygenic score adds thousands of tiny genetic nudges into one number. Why that number says a lot about your height and almost nothing about coriander, and how to read a disease percentile without being scared by it.
Most things your DNA report tells you rest on one variant. Lactose tolerance, earwax, how fast you clear caffeine: one position, a clear answer.
Most of what people actually wonder about does not work that way. Height, appetite, whether you are a morning person, your odds of type 2 diabetes: each is shaped by thousands of variants, every one of them a nudge too small to notice alone. A polygenic score adds those nudges up.
It is a genuinely useful idea. It is also the easiest result in genetics to misread, in both directions.
What a polygenic score actually is
Large studies compare millions of genetic variants between people who differ in a trait. Each variant gets a weight: how much one copy of it moves the trait, on average. Your score is simply the sum, over all of those variants, of how many copies you carry times the weight.
The PGS Catalog collects thousands of these published scores, each with the study it came from. Where no published score exists, one can be built from a study's full results by keeping the strongest, independent signals.
The number means nothing on its own
A raw score is a figure like 4.37. What it means only shows up when you compare it with other people.
The honest way to do that is to compute the same score for a real reference group and see where you land. The 1000 Genomes Project published whole genomes for 2,504 people from 26 populations, grouped into five: African, American, East Asian, European and South Asian. If your score is higher than 88 of every 100 of them, you are at the 88th percentile.
Two things follow from that:
- A percentile is relative. It says where you sit in a group, not what will happen to you.
- The group matters. People are best compared with the reference group closest to their ancestry, and 1000 Genomes has no Middle Eastern or North African group. For readers with that ancestry, Europe is the nearest comparison, and it should say so.
How much genes actually decide
This is the part most reports leave out, and it changes everything.
Approximate, in people of European ancestry. Height from Yengo et al. (2022). Morning person and coriander are the in-sample ceilings of the scores HealthOS uses, which flatter them, and the real figure is lower.
A height score explains around 40% of the differences in height between people of European ancestry. That is a lot. Someone at the 90th percentile of a good height score has roughly an 85 in 100 chance of being taller than average.
A score for liking coriander explains under 1%. Even someone in the top 1% of that score has only about a 55 in 100 chance of liking coriander more than average, barely better than a coin flip. What you grew up eating matters far more than any of these variants.
Both are real polygenic scores, computed the same way. One is a strong prediction and the other is close to noise. Any report that words them with the same confidence is misleading you, and "you are very likely to love coriander" is the kind of sentence to distrust.
The question to ask of any polygenic result is not "what is my percentile?" but "how much does this score explain?"
Disease scores: relative risk is not your risk
For conditions like coronary heart disease or type 2 diabetes, a high percentile sounds alarming. "Higher than 88 in 100 people" is true, and it still does not say how likely you are to get the condition.
A validated score comes with a measured strength: how much the odds rise for each step up the scale. For the type 2 diabetes score HealthOS uses, independent validations put that at about 1.2 times per standard deviation. In the top 10%, that works out to about 1.25 times the average person's odds.
To turn that into something you can picture, you need the baseline. As an illustration: if 10 in 100 people like you develop a condition, 1.25 times that is about 12 or 13 in 100. That is a real difference, and not a verdict. Weight, activity and diet move type 2 diabetes risk far more than these genes do.
Some scores are much stronger. In the study that put polygenic scores on the map, the top 8% of a coronary heart disease score carried about three times the average risk, comparable to some single-gene conditions. Most scores are weaker than that, and researchers have long warned that they sort populations far better than they predict for any one person.
Why your ancestry changes how much to trust it
Almost all of the large genetic studies behind these scores were done in people of European ancestry. Scores built from them work best in Europeans. One major analysis found prediction accuracy around twice as low in East Asian ancestry and several times lower in African ancestry.
That is not a reason to ignore your result. It is a reason for the result to tell you which group it was built in, and to be more cautious the further you are from it.
A 23andMe file versus a whole genome
A consumer genotyping chip, such as 23andMe, AncestryDNA or MyHeritage, reads around 600,000 positions. Measured against the scores HealthOS uses, a real chip carries about a third of the variants they need. The rest are not zeroes, they are unknowns, so a score summed over only what the chip reads is a different number, not a rough version of the right one.
The fix is imputation. Stretches of DNA are inherited in blocks, so the positions a chip does read reveal which common blocks you carry, and a reference panel fills in the rest. In our tests on a genome held out of the reference, the filled-in genotypes matched the real ones with a correlation (r²) of 0.97 at common positions. That was a European genome; accuracy is lower for ancestries the reference panel represents poorly.
A whole-genome file, from Nebula Genomics, Dante Labs or similar, needs none of this. It lists every position where you differ from the reference genome, so everything else is known.
How to read your own
- Check how much the score explains before reading anything into the percentile.
- Treat a percentile as a tendency within a group, never a diagnosis and never destiny.
- For a disease score, ask for the relative risk and a baseline, not just the percentile.
- Check which ancestry the score was built in, and how close yours is.
- Act on what you can measure. For type 2 diabetes, an HbA1c blood test tells you more about where you stand today than any genetic score.
How HealthOS handles this
The free DNA report at healthosx.com/dna sticks to well-established single-variant findings. If you store your file, Ask your DNA can also read about 320 polygenic scores, from food liking and sleep to fitness and disease tendencies. It follows the rules on this page:
- Tastes and traits come back in plain words, and only as confident as the genetics allow. A strong height score can say "very likely", while coriander usually gets "your genes don't push either way".
- Health tendencies come with numbers, such as "higher than 45 in 100 people with ancestry like yours", with a relative risk wherever the score's strength has been measured independently.
- Every result is placed among the 2,504 people of 1000 Genomes, in the group closest to your ancestry, and says which group that is.
- Chip files are filled in first. Imputation runs in about ten minutes after you store the file, and whole-genome files are scored right away.
Read your own file free, then ask it anything. healthosx.com/dna
References
- 1.Lambert et al., "The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation," Nature Genetics (2021)
- 2.Khera et al., "Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations," Nature Genetics (2018)
- 3.Martin et al., "Clinical use of current polygenic risk scores may exacerbate health disparities," Nature Genetics (2019)
- 4.Yengo et al., "A saturated map of common genetic variants associated with human height," Nature (2022)
- 5.May-Wilson et al., "Large-scale GWAS of food liking reveals genetic determinants and genetic correlations with distinct neurophysiological traits," Nature Communications (2022)
- 6.The 1000 Genomes Project Consortium, "A global reference for human genetic variation," Nature (2015)
- 7.Browning et al., "A one-penny imputed genome from next-generation reference panels," American Journal of Human Genetics (2018)
- 8.Wald and Old, "The illusion of polygenic disease risk prediction," Genetics in Medicine (2019)
For informational purposes only. For medical advice or diagnosis, consult a professional.
Written by
HealthOS Research