arXiv:cs.LG(机器学习,全量分类)· Yuan Huang, Zirui Song, Xiuying Chen·· 15 小时前AI 评分46
MLLM 评委真的在评判编辑效果吗?EditJudgeBias 审计图像编辑评估中的偏差
Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation
AI 导读
研究者提出 EditJudgeBias 反事实基准,包含 1,196 个真实编辑样本和跨四个评估位点注入的 13 种线索,并用校准的多模态验证器、对照与人工检查确认编辑质量未变。对五个 MLLM 评委的审计显示,保质量的线索即可让所有评委的评分偏移超出自身噪声下限,伪造的多数意见会抬高评分,交换候选顺序最多逆转 60.9% 的成对判断,编辑区域线索还会降低与人类判断的一致性。
正文
Abstract:Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may themselves alter the quality being evaluated. A judgment shift can therefore be attributed to bias only when the intervention is verified to preserve the underlying editing quality. To address this challenge, we introduce EditJudgeBias, a counterfactual benchmark with verified quality preservation, comprising 1,196 real editing samples and 13 cues injected across four evaluation sites. We verify quality preservation for the requested edit using calibrated multimodal validators, controls, and human inspection. We then audit five MLLM judges along three complementary dimensions: invariance to quality-preserving cues, agreement with human judgments, and stability of pairwise preferences. Importantly, observed shifts are evaluated against each judge's own zero-dose and re-query noise floors rather than against zero. Experiments show that quality-preserving cues move every judge beyond its own noise. Fabricated majority opinions increase ratings, irrelevant visual elements cause larger shifts than whole-image manipulations, and swapping candidate order reverses up to 60.9% of pairwise decisions. Edit-region cues also tend to reduce human agreement. The three measures characterize judges differently, showing that robustness cannot be captured by a single metric.
| Comments: | 30 pages, 9 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01670 [cs.CV] |
| (or arXiv:2610.01670v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01670 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuan Huang [view email]
[v1]
Thu, 1 Oct 2026 13:30:07 UTC (2,532 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org