TDG Programme Blog DH Concierge Book a Call
Future of Medicine

They Built an AI to Predict 1,000 Diseases. HbA1c Still Beat It.

Two of the most ambitious projects in predictive medicine have published in the last two years. One forecasts more than a thousand diseases decades ahead. The other predicts over three thousand from blood biomarkers alone. What they found, taken together, is more interesting than either headline — and less flattering to everybody than it first appears.

STEPHEN DUNCAN FDN-P BSC HONS MSC · DETECTIVE HEALTH · AUGUST 2026
Large datasets · genuine advance · prediction is not treatment

I want to write this carefully, because there is a lazy version of it that would be very satisfying and not quite honest.

The lazy version says: the AI labs have spent hundreds of millions arriving at what functional medicine has been saying for thirty years. Measure early. Track trajectories. Treat individuals rather than averages. Told you so.

There is something in that. There is also rather less than the people who would enjoy saying it might hope, and the interesting findings sit in the gap.

Delphi-2M

Nature, 2025

A team led by Moritz Gerstung at the German Cancer Research Center and Ewan Birney at EMBL-EBI adapted a GPT-style transformer — the same architecture behind conversational AI — to predict diseases rather than words.

Each diagnosis, lifestyle factor and demographic marker becomes a token. The model learns how they follow one another over a lifetime, then predicts what comes next and roughly when.

Trained on 400,000 UK Biobank participants, covering 1,256 disease codes, and externally validated on 1.9 million people in the Danish national registry. Average AUC 0.76 for the next diagnosis, 0.73 for diagnoses a year ahead, and 0.67 in Denmark with no retraining at all. For many conditions it matched purpose-built single-disease tools like QRisk.

That last detail is the impressive one. A single general model, transferring across health systems and countries without adjustment, performing comparably to tools built for one disease each.

MILTON, And The Finding Nobody Reported

The second project comes from AstraZeneca. MILTON is a machine-learning framework using blood biomarkers to predict 3,213 diseases in UK Biobank participants, identifying incident cases that were undiagnosed at recruitment.

Here is the part worth sitting with: MILTON, working from ordinary blood biomarkers, largely outperformed available polygenic risk scores.

Your bloods predicted your future health better than your genome did. In a dataset of half a million people, using the best genetic risk scores available.

This is the finding I would put in front of anyone who has had a genetic test and come away frightened. Genotype is a starting condition. Biochemistry is a running measurement of what that starting condition is currently doing, in your body, in your circumstances — and it turns out to carry more predictive information.

And Then The Detail In The Nature Paper

Delphi-2M outperformed a multi-disease predictor trained on 67 UK Biobank biomarkers. It is a genuinely strong model.

But for diabetes, it lost

The paper reports that Delphi-2M’s predictive power trailed some strong lab markers — specifically HbA1c for diabetes.

A transformer model, trained on four hundred thousand people, validated across two national health systems, forecasting more than a thousand conditions — and for one of the most consequential diseases in the developed world, it was beaten by a single blood test that costs a few pounds and has been available since the 1970s.

I find that genuinely delightful, and I want to be careful about what it does and does not mean.

It does not mean the AI work is overblown. Predicting a thousand diseases at once is a different and harder problem than predicting one well, and the model was never claiming to beat every specialist tool.

What it means is narrower and more useful: when a biomarker sits close to the mechanism of a disease, it is very hard to beat. HbA1c is not correlated with glycaemic dysfunction. It is glycaemic dysfunction, integrated over three months. No amount of pattern-matching over health records improves on measuring the thing itself.

What Is Actually Being Vindicated

Not functional medicine. A principle that functional medicine happens to share with a good deal else.

That disease has a trajectory, and the trajectory is measurable before the diagnosis. Both projects rest entirely on this. Delphi-2M works because health histories contain signal about what comes next. MILTON works because biochemistry drifts before it breaks.

And that the individual measurement carries more information than the population average. This is the whole point of both models — not what happens to people like you, but what your own data says about you.

Neither of those is a functional medicine idea. They are ideas about biology that functional medicine has organised itself around, sometimes rigorously and sometimes not.

Which is where I have to be honest

Being early to a correct principle is not the same as being right about its applications. My own field has been early to trajectory-based thinking and simultaneously wrong about a great deal else — and I have published a running list of the things I taught before checking properly, including a supplement claim that does not survive its own trial data, an in-range thyroid number treated as a diagnosis, and a movement position I held for thirty-seven years and had to withdraw.

So the honest claim is modest. The direction of travel supports the principle. It does not retrospectively validate everything done in its name, mine included. Anyone in this field reading the Nature paper as vindication should read their own back catalogue first.

What These Models Cannot Do

The same limitation that applies to the Alzheimer’s blood test applies here, and it is the reason this matters clinically rather than academically.

Prediction is not intervention. Delphi-2M can tell a population that a given health history carries an elevated probability of a given diagnosis. It cannot tell you what to do about it, whether doing anything helps, or which of the modifiable inputs in that person matters most.

And its own authors are clear about the limits: UK Biobank captures only a person’s first encounter with each disease, recruitment is biased toward healthier and wealthier participants, and performance dropped when moved to Danish data. An AUC of 0.76 is a useful model. It is not a crystal ball, and the gap between those two things is where most of the hype lives.

The difference between what they do and what I do

They work at population scale with enormous datasets and no individual. Four hundred thousand people, statistical power I will never have, and no ability to ask any one of them what changed last year.

I work with one person and a fraction of the data. Which means far weaker inference and far better context — I know about the bereavement, the new job, the medication, the thing they did not put on the form.

Neither replaces the other. But it is worth being clear that a clinician's pattern recognition is not a small version of a transformer model. It is a different epistemic activity, with different failure modes, and mine fails through confirmation bias in ways a model does not.

What I Take From It

Blood biomarkers beat polygenic risk scores. That is the single most useful finding here for anyone anxious about a genetic result, and it points spending toward the thing that is both cheaper and changeable.

Proximity to mechanism beats sophistication. HbA1c beating a transformer for diabetes is a lesson about test selection, not about AI. Measure the thing, close to where it happens, rather than inferring it from further away.

And the modifiable layer is still nobody’s department. These models predict. The National Health Service treats what has arrived. The space between — the years where something is drifting and could be moved — remains largely unoccupied, which is the space I work in and the reason the work exists.

If the AI labs eventually occupy it properly, that will be a good outcome and I will say so. They are not there yet.

Educational content, not medical advice. The models described are research tools, not clinical services, and are not available to order. Nothing here is a reason to seek predictive testing outside a clinical pathway.

Related

They Built an AI to Predict 1,000 Diseases The Alzheimer’s Blood Test Is Real Nobody Has a Digital Twin More from the blog →

Test, don’t guess

The trajectory is measurable years before the diagnosis. That is the whole argument, and it is now other people’s argument too.

See testing options Ask AIdan →