arXiv:cs.LG· Baoteng Li, Wenzhuo Wu, Kongming Liang, Zhanyu Ma·· 4 小时前AI 评分33
面向多主体图像生成的 Visual Jev 参考绑定奖励
Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
AI 导读
研究提出参考绑定 Visual Jev 奖励,仅在指定条件与对应参考身份同时成立时给主体相关问题打正标签,用于多主体图像生成的训练信号。
正文
Abstract:Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09328 [cs.CV] |
| (or arXiv:2610.09328v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09328 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Baoteng Li [view email]
[v1]
Wed, 7 Oct 2026 02:34:09 UTC (488 KB)
来源:arXiv:cs.LG · arxiv.org