Researchers trained two language models (LLaMA-8B-Instruct and ClinicalModernBERT) on 9,263 narrative death investigation reports from the New York City Office of Chief Medical Examiner for 2022-2023, using a two-step pipeline that first extracted 18 expert-defined textual indicators of overdose. LLaMA-8B-Instruct reached 88% positive predictive value, 92% sensitivity and 76% specificity (macro-F1 0.85), and ClinicalModernBERT reached 83% PPV and 67% specificity. Both exceeded the office's existing surveillance tool and conventional logistic regression on the same reports, in an internal evaluation.
Why it is interesting: Tests whether language models can read unstructured death investigation text well enough for cause-of-death surveillance, with specificity the weaker measure.