Researchers built LLM systems combining guideline-enriched prompting, supervised fine-tuning and a self-correction step to identify in-hospital cardiac arrest events and their locations from electronic health record notes, benchmarked against physician-adjudicated chart review. On a curated corpus the fine-tuned model reached precision 0.88, recall 1.00 and F1 0.94, against 0.68, 0.91 and 0.78 for an ICD-code algorithm. In an unselected cohort of 45,525 encounters the fine-tuned model underperformed, while a prompting-plus-self-correction configuration achieved precision 0.79.
Why it is interesting: The gap between performance on a curated corpus and in an unselected cohort of 45,525 encounters, where the fine-tuned model underperformed.