COIOS
Weighing data and models Item new peer-reviewed

An LLM identified NYC overdose deaths from investigator narratives with 92% sensitivity

Two language models trained on 9,263 New York City death investigation narratives from 2022-2023 identified suspected overdose deaths with 92% sensitivity and 88% positive predictive value, outperforming the medical examiner's existing surveillance tool and logistic regression.

Researchers trained two language models (LLaMA-8B-Instruct and ClinicalModernBERT) on 9,263 narrative death investigation reports from the New York City Office of Chief Medical Examiner for 2022-2023, using a two-step pipeline that first extracted 18 expert-defined textual indicators of overdose. LLaMA-8B-Instruct reached 88% positive predictive value, 92% sensitivity and 76% specificity (macro-F1 0.85), and ClinicalModernBERT reached 83% PPV and 67% specificity. Both exceeded the office's existing surveillance tool and conventional logistic regression on the same reports, in an internal evaluation.

Why it is interesting: Tests whether language models can read unstructured death investigation text well enough for cause-of-death surveillance, with specificity the weaker measure.

Source
American journal of public health, 1 October 2026
DOI
10.2105/ajph.2026.308698
Type
Peer-reviewed article
Design
Model development and internal evaluation on 9,263 narrative death investigation reports, New York City OCME, 2022–2023, two LLM families against the incumbent tool and logistic regression
Verdict
New finding
Driver
AI in medicine
Driver
Alcohol and drug use