跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Luyao Tang, Cheng Chen·· 14 小时前AI 评分44

OmniMed-Jev:用决策原生接口校准医疗 LVLM 置信度

OmniMed-Jev: Calibrating LVLM Confidence for Trustworthy Medical Multimodal Decisions via System One

AI 导读

OmniMed-Jev 将医疗决策表示为对运行时候选集的 Choice、Noul 或 Score 决策并输出完整概率分布,在与同骨干、同数据、同训练的生成式基线对比中,校准误差最多降低一个数量级、可靠性误差最多降低两倍,点预测性能相当,仅计数任务生成式基线仍占优。该接口统一支持多种影像模态与预测任务,使异构输出变为可比较的概率;作者强调结果不构成临床可用性证据,代码已公开。

正文

View PDF HTML (experimental)

Abstract:Medical models are judged not only on correctness, but on whether reported confidence matches actual accuracy. Generalist multimodal medical models have expanded what a single model can perceive, yet they still express bounded decisions such as diagnoses, findings or cell counts as generated text, so the reported probability reflects the next token rather than the decision itself. Motivated by decision-native interfaces such as Jev, we introduce OmniMed-Jev, which represents each medical decision as a Choice, Noul or Score decision over a runtime-supplied candidate set and returns a full distribution over that set: mutually exclusive classes, binary presence of a finding, or a bounded ordered value. The design is omni in three respects: it accepts diverse imaging modalities, covers different prediction tasks, and expresses them through one candidate-conditioned probability model, so heterogeneous outputs become comparable probabilities rather than task-specific strings. In an interface-controlled comparison against a generative baseline trained on the same backbone, data and schedule, OmniMed-Jev's reported probabilities track observed correctness far more closely, reducing calibration error by up to an order of magnitude and reliability error by up to two, while point-prediction performance remains comparable; counting is the one family where the generative baseline stays ahead. Making the decision distribution the model's output is not a format change but what turns reported numbers into probabilities that mean what they say. These results support explicit decision modeling as a way to make reported confidence meaningful within the evaluated tasks, and they are not evidence of clinical readiness: the comparison cannot separate the interface from associated training differences, which we state alongside the results. Code is available at this http URL.
Comments: We introduce OmniMed-Jev, a decision-native interface based on Jev that outputs Choice, Noul or Score decisions with calibrated probabilities
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00381 [cs.LG]
  (or arXiv:2610.00381v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00381

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Luyao Tang [view email]
[v1] Wed, 30 Sep 2026 08:53:54 UTC (117 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org