跳到正文
arXiv:cs.CL· Weiwei Ma, Xiaobing Yu, Peijie Qiu, Jin Yang, Zhaoqi An, Xuanzhao Dong, Xiaoqi Zhao, Xiaofeng Liu·· 3 小时前

ActiveMedAgent:面向多模态医疗诊断的成本感知轨迹学习

ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis

AI 导读

研究提出 ActiveMedAgent 框架,在冻结的 API 视觉语言模型之上,用概率分布追踪候选诊断并按诊断效用减成本为每次检查打分,再离线训练轻量 MLP 控制器学习何时补充证据、何时给出结论。在三个常用基准上,基于轨迹的策略学习均优于无引导采集和全模态基线;175 个案例中智能体用更少通道得出正确诊断而全模态基线失败,显示信息过载效应。该工作已被 EMNLP 2026 收录。

正文

View PDF HTML (experimental)

Abstract:Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.
Comments: EMNLP 2026
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as: arXiv:2610.11140 [cs.LG]
  (or arXiv:2610.11140v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11140

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xiaofeng Liu [view email]
[v1] Thu, 8 Oct 2026 03:06:11 UTC (459 KB)

来源:arXiv:cs.CL · arxiv.org