arXiv:cs.LG· Omar Elfatairy, Maria A. Bravo, Jessica Bader, Zeynep Akata·· 3 小时前AI 评分55
NegT2IBench:面向文生图模型否定约束的极性基准
NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models
AI 导读
研究者发布 NegT2IBench,一个包含 4,800 条提示词的文生图否定约束基准,按极性组织提示词以分离否定效应与提示词复杂度。在 11 个 T2I 模型和 211,200 张图像上,9 个模型在单条否定语句上的得分低于单条肯定语句,41.5% 的失败语句恰好渲染了提示词禁止的内容;其检测器评分与人类标注的一致性接近大 30 倍的视觉语言裁判,而 GPU 内存占用仅为其一小部分。
正文
Abstract:Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute types and four relation categories. Prompts are organized by polarity: the number of positive statements that must hold and negated statements that must not, each ranging from 0 to 2. Varying the two independently separates the effect of negation from the effect of prompt complexity. Our detector-based scoring is reproducible, auditable, and pinpoints which requirement failed. On 600 images with three-annotator labels, it agrees with humans as closely as vision-language judges up to 30x larger, while using only a fraction of their GPU memory. Across eleven T2I models and 211,200 images, nine score lower on a single negated statement than on a single positive one. Per-statement scoring reveals that the loss is largest for color and near zero for proximity, and that 41.5% of failed statements render exactly what the prompt forbids. Rendering what a prompt asks for and withholding what it forbids are distinct capabilities that an aggregate compositional score cannot distinguish. NegT2IBench measures the latter directly, providing a controlled testbed for diagnosing negation failures and developing methods to overcome them.
| Comments: | *Equal contribution |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03084 [cs.CV] |
| (or arXiv:2610.03084v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03084 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Maria A. Bravo [view email]
[v1]
Fri, 2 Oct 2026 10:03:56 UTC (15,814 KB)
来源:arXiv:cs.LG · arxiv.org