arXiv:cs.AI· Shuo Yang, Lihao Fang, Yi Zhang, Haixiang Wang, Xincheng Ye, Shufan Chen, Jipeng Guo, Youqing Wang·· 4 小时前AI 评分37
重新思考少步扩散 Transformer 中缓存什么:求解器感知的目标选择
Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection
AI 导读
针对少步蒸馏后相邻采样步间隔变大、复用张量误差升高的问题,研究者提出 AutoTarget,可根据模型、求解器和复用调度自动选择误差最低的缓存张量。该方法在蒸馏图像与视频 DiT 上减少了 DiT 前向评估次数和缓存存储占用,生成质量接近未缓存版本;在 PixArt-LCM 和 FLUX.1-schnell 上,其校准排序与留出缓存运行结果一致,核心实现已在 GitHub 开源。
正文
Abstract:Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tensor computed at an earlier step. Most caching methods decide in advance which tensor to reuse. After distillation, adjacent sampling steps are farther apart. Reusing a tensor across this larger gap introduces more error, so choosing what to cache becomes especially important. We therefore introduce AutoTarget, a method that chooses the cached tensor for a given model, solver, and reuse schedule. AutoTarget uses a small set of runs without cache reuse to measure the error caused by reusing each candidate tensor, then selects the candidate with the lowest error. We also analyze how an error at one reuse step affects the final sample. For Euler sampling, we identify cache targets that produce the same trajectory and show why a stored solver update may not. Experiments on distilled image and video DiTs show that the best cache target changes with the model, image resolution, and solver. AutoTarget reduces DiT evaluations and retained cache storage. Generation quality remains close to the corresponding uncached run. On the tested PixArt-LCM and FLUX.1-schnell settings, its calibration ranking matches the ranking from held-out cached runs. To help others reproduce the method, we provide its core implementation on GitHub at this https URL.
| Comments: | 20 pages, 9 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03577 [cs.CV] |
| (or arXiv:2610.03577v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03577 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shuo Yang [view email]
[v1]
Fri, 2 Oct 2026 16:51:22 UTC (7,847 KB)
来源:arXiv:cs.AI · arxiv.org