arXiv:cs.CL· Weiwei Ma, Xiaobing Yu, Peijie Qiu, Jin Yang, Zhaoqi An, Xuanzhao Dong, Xiaoqi Zhao, Xiaofeng Liu·· 3 小时前
ActiveMedAgent:面向多模态医疗诊断的成本感知轨迹学习
ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis
AI 导读
研究提出 ActiveMedAgent 框架,在冻结的 API 视觉语言模型之上,用概率分布追踪候选诊断并按诊断效用减成本为每次检查打分,再离线训练轻量 MLP 控制器学习何时补充证据、何时给出结论。在三个常用基准上,基于轨迹的策略学习均优于无引导采集和全模态基线;175 个案例中智能体用更少通道得出正确诊断而全模态基线失败,显示信息过载效应。该工作已被 EMNLP 2026 收录。
正文
Abstract:Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.
| Comments: | EMNLP 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM) |
| Cite as: | arXiv:2610.11140 [cs.LG] |
| (or arXiv:2610.11140v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11140 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaofeng Liu [view email]
[v1]
Thu, 8 Oct 2026 03:06:11 UTC (459 KB)
来源:arXiv:cs.CL · arxiv.org