跳到正文
arXiv:cs.LG· Wanning He, Yuyao Zhang, Yu-Wing Tai·· 4 小时前AI 评分42

RefRoute:通过紧凑残差条件与空间路由将条件成本与参考解耦

RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing

AI 导读

RefRoute 框架通过紧凑残差条件与空间路由两种机制,解决多参考图像生成中参考 token 数量与注意力开销随参考数量和分辨率增长的问题。

正文

View PDF HTML (experimental)

Abstract:Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve $18.3\times$ and $14.2\times$ speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
Comments: 19 pages. Wanning He and Yuyao Zhang contributed equally and share first authorship
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.07720 [cs.CV]
  (or arXiv:2610.07720v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.07720

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wanning He [view email]
[v1] Tue, 6 Oct 2026 04:16:27 UTC (27,488 KB)

来源:arXiv:cs.LG · arxiv.org