arXiv:cs.AI· Chufan Shi, Cheng Yang, Tiannuo Yang, Isadora White, Yiwei Chen, Taylor Berg-Kirkpatrick, Xuezhe Ma·· 6 小时前AI 评分53
统一多模态模型视觉拒绝研究:提出 Draw-or-Decline 基准与 VisTA 训练方法
Visual Abstention in Unified Multimodal Models
AI 导读
arXiv 论文提出视觉拒绝(visual abstention)概念:当视觉变换请求不可行时,模型应识别并拒绝生成。
正文
Abstract:Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.
| Comments: | 25 pages, 6 figures, 13 tables. Project page: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.07887 [cs.CL] |
| (or arXiv:2610.07887v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07887 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chufan Shi [view email]
[v1]
Tue, 6 Oct 2026 07:33:42 UTC (1,674 KB)
来源:arXiv:cs.AI · arxiv.org