COIOS
Weighing data and models Item new preprint

Fine-tuned LLM beat ICD codes for in-hospital cardiac arrest on a curated corpus

Large language models reading clinical notes identified in-hospital cardiac arrest with precision 0.88 and recall 1.00 on a curated corpus against 0.68 and 0.91 for ICD codes, though a prompting configuration reached precision 0.79 in an unselected cohort of 45,525 encounters.

Researchers built LLM systems combining guideline-enriched prompting, supervised fine-tuning and a self-correction step to identify in-hospital cardiac arrest events and their locations from electronic health record notes, benchmarked against physician-adjudicated chart review. On a curated corpus the fine-tuned model reached precision 0.88, recall 1.00 and F1 0.94, against 0.68, 0.91 and 0.78 for an ICD-code algorithm. In an unselected cohort of 45,525 encounters the fine-tuned model underperformed, while a prompting-plus-self-correction configuration achieved precision 0.79.

Why it is interesting: The gap between performance on a curated corpus and in an unselected cohort of 45,525 encounters, where the fine-tuned model underperformed.

Source
medRxiv, 27 September 2026
DOI
10.64898/2026.09.24.26363973
Type
Preprint
Design
Retrospective LLM case-ascertainment study against physician-adjudicated chart review; curated evaluation corpus plus 45,525-encounter unselected validation cohort
Verdict
New finding
Driver
AI in medicine
Driver
Heart and stroke