跳到正文
HuggingFace Daily Papers·· 2 天前

Iris-3B:像素空间扩散的预训练、转换与微调,超越潜空间?

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

AI 导读

研究团队从头预训练了 3B 参数的像素空间文本到图像 Transformer Iris-3B,采用 $256\to512\to1024$ 课程,并把潜空间模型 FLUX.2 Klein base 4B 转换为像素空间。

正文

View PDF HTML (experimental)

Abstract:Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a $256\to512\to1024$ curriculum, after first ablating the prediction target and representation alignment at $256^2$ to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on $4\times$ DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at $1024^2$. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.
Comments: 19 pages, 13 figures, 6 tables. Project lead: Hanqiu Li Cai. Code and models: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
ACM classes: I.2.10; I.4.4
Cite as: arXiv:2610.09450 [cs.CV]
  (or arXiv:2610.09450v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.09450

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chema Garabito [view email]
[v1] Wed, 7 Oct 2026 05:07:09 UTC (18,598 KB)

来源:HuggingFace Daily Papers · arxiv.org