跳到正文
arXiv:cs.LG· Kihyuk Yoon, Lingchao Mao, Catherine Chong, Todd J. Schwedt, Chia-Chun Chiang, Jing Li·· 3 小时前

PaReGTA:面向 EHR 分析的时间感知 LLM 患者表征框架

PaReGTA: A Temporally Aware LLM-Based Patient Representation Framework for EHR Analytics

AI 导读

PaReGTA 是一个基于 LLM 的编码框架,将纵向 EHR 事件转为带显式时间线索的就诊级模板文本,经轻量对比微调学习就诊嵌入向量,并用混合时间池化聚合为固定维度患者表征。在 All of Us 研究计划 39,088 名偏头痛患者队列上,PaReGTA-Gap + LightGBM 在 7,818 人固定留出测试集中取得的 AUC、准确率和 F1 为所评估方法中最高。

正文

View PDF HTML (experimental)

Abstract:Temporal information in structured electronic health records (EHRs) is often lost in sparse one-hot or count-based representations, while sequence models can be costly and data-hungry. We propose PaReGTA, an LLM-based encoding framework that (i) converts longitudinal EHR events into visit-level templated text with explicit temporal cues, (ii) learns domain-adapted visit embeddings via lightweight contrastive fine-tuning of a sentence-embedding model, and (iii) aggregates visit embeddings into a fixed-dimensional patient representation using hybrid temporal pooling that captures both recency and globally informative visits. The resulting fixed-dimensional patient representations can be used with conventional downstream machine-learning models. To examine factor-level sensitivity, we use PaReGTA-RSS (Representation Shift Score), a prespecified factor-removal analysis that recomputes patient representations after removing clinically defined factor groups and quantifies the resulting change in the fitted logit of a fixed logistic-regression model. We evaluated PaReGTA in a cohort of 39,088 patients with migraine from the All of Us Research Program (AoU) on a retrospective patient-level classification task with an EHR-derived target. In an exploratory comparison on the fixed held-out test cohort of 7,818 patients, PaReGTA-Gap + LightGBM, used as a post hoc analytical reference, had the highest reported AUC, accuracy, and F1 values among the evaluated sparse, BERT-based EHR, and recurrent approaches in this cohort.
Comments: 37 pages, 5 figures, 21 tables
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2602.19661 [cs.LG]
  (or arXiv:2602.19661v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2602.19661

arXiv-issued DOI via DataCite

Submission history

From: Kihyuk Yoon [view email]
[v1] Mon, 23 Feb 2026 10:09:50 UTC (1,038 KB)
[v2] Sun, 1 Mar 2026 13:04:03 UTC (1,038 KB)
[v3] Thu, 8 Oct 2026 14:24:16 UTC (1,428 KB)

来源:arXiv:cs.LG · arxiv.org