The authors trained an autoregressive generative model on longitudinal electronic health records from millions of patients, explicitly representing the irregular time gaps between encounters, then adapted it with parameter-efficient supervised fine-tuning for pan-cancer risk stratification. Across five large EHR cohorts, the adapted model improved prediction of a first cancer diagnosis within a five-year window relative to the foundational representation alone; no absolute discrimination figures are given in the abstract. The authors frame the work as retrospective evidence supporting prospective evaluation for prioritising patients for risk-based screening, including pancreatic and ovarian cancer.
Why it is interesting: Tests whether a general-purpose generative model of patient histories, rather than a task-specific risk equation, can identify who to screen, so far only retrospectively.