跳到正文
arXiv:cs.LG· Omar Elfatairy, Maria A. Bravo, Jessica Bader, Zeynep Akata·· 3 小时前AI 评分55

NegT2IBench:面向文生图模型否定约束的极性基准

NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models

AI 导读

研究者发布 NegT2IBench,一个包含 4,800 条提示词的文生图否定约束基准,按极性组织提示词以分离否定效应与提示词复杂度。在 11 个 T2I 模型和 211,200 张图像上,9 个模型在单条否定语句上的得分低于单条肯定语句,41.5% 的失败语句恰好渲染了提示词禁止的内容;其检测器评分与人类标注的一致性接近大 30 倍的视觉语言裁判,而 GPU 内存占用仅为其一小部分。

正文

View PDF HTML (experimental)

Abstract:Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute types and four relation categories. Prompts are organized by polarity: the number of positive statements that must hold and negated statements that must not, each ranging from 0 to 2. Varying the two independently separates the effect of negation from the effect of prompt complexity. Our detector-based scoring is reproducible, auditable, and pinpoints which requirement failed. On 600 images with three-annotator labels, it agrees with humans as closely as vision-language judges up to 30x larger, while using only a fraction of their GPU memory. Across eleven T2I models and 211,200 images, nine score lower on a single negated statement than on a single positive one. Per-statement scoring reveals that the loss is largest for color and near zero for proximity, and that 41.5% of failed statements render exactly what the prompt forbids. Rendering what a prompt asks for and withholding what it forbids are distinct capabilities that an aggregate compositional score cannot distinguish. NegT2IBench measures the latter directly, providing a controlled testbed for diagnosing negation failures and developing methods to overcome them.
Comments: *Equal contribution
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.03084 [cs.CV]
  (or arXiv:2610.03084v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.03084

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Maria A. Bravo [view email]
[v1] Fri, 2 Oct 2026 10:03:56 UTC (15,814 KB)

来源:arXiv:cs.LG · arxiv.org