arXiv:cs.LG· Arshia Hemmat, Amirhossein Vahidi, Amitis Shidani, Mohammad Vali Sanian, Hesam Asadollahzadeh, Aryan Yazdan Parast, Mohammad Lotfollahi·· 4 小时前AI 评分42
ORCA:定位文生图扩散模型的组合性失败
ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
AI 导读
一种名为 ORCA(Orthogonal Residual Compositional Alignment)的方法通过单一辅助损失,将扩散 Transformer 的潜变量与冻结视觉编码器的低秩目标对齐,在 DiT-B/2、DiT-L/2、U-ViT-L 三个骨干上均以零推理成本提升 FID 与 GenEval。
正文
Abstract:Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.
| Comments: | Accepted at NeurIPS 2026. 26 pages, 4 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09841 [cs.CV] |
| (or arXiv:2610.09841v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09841 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Arshia Hemmat [view email]
[v1]
Wed, 7 Oct 2026 11:03:58 UTC (1,565 KB)
来源:arXiv:cs.LG · arxiv.org