COIOS
Emergent technology emerging strengthening

AI in medicine

Whether machine learning predicts, detects and decides better than the tools we have, and whether that has yet changed an outcome for a population.

Where the evidence stands

Papers claiming that a machine-learning model predicts death, or a disease, better than the standard tool now appear weekly, and a majority are trained and tested in a single dataset with no external validation. Where validation has been done properly, the pattern is consistent: modest gains in discrimination over well-built conventional models, larger gains where the conventional model was poor or the data unusually rich — full electronic records, imaging, wearables — and almost no evidence that better prediction has changed an outcome, because prediction and action are different problems.

Detection is further along than prediction. Reading images is what the technology does best, and the first randomised trials of AI-supported screening — mammography in Sweden most prominently — show more cancers found at similar false-positive rates with less radiologist time; whether that becomes fewer interval cancers or fewer deaths is the question the same trials are now following. Language models have entered clinical work as note-takers and document summarisers faster than any regulator has assessed them, on the strength of time saved rather than outcomes measured.

At population scale the picture is thinner still: national statistical offices have begun testing machine-learning methods for coding causes of death, imputation and nowcasting, with mixed published results. The honest state is that the technology is real, the claims outrun the evidence, and the studies that would settle it — external validation, impact on decisions, performance across populations — are rare.

Status
emerging
Direction
strengthening
Would change our view
An externally validated model showing a large, reproducible gain over the best conventional model on a population-level outcome; a screening trial in which AI support reduced interval cancers or mortality; or a trial in which acting on a model's predictions changed mortality.
Trackers
Gaps
Validation across countries and ancestries; comparisons against strong rather than weak baselines; outcome evidence for language-model tools; anything from the offices on operational use.

At a glance

Status
emerging
Direction
strengthening
Last changed
Evidence
Countries

Related drivers

Others we follow in Emergent technology.