arXiv:cs.LG· Huawei Lin, Tony Geng, Zhaozhuo Xu, Weijie Zhao·· 5 小时前AI 评分39
VTBench:评估自回归图像生成中的视觉分词器
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
AI 导读
研究者推出 VTBench,一个系统评估视觉分词器(VT)的基准,覆盖图像重建、细节保留和文本保留三项核心任务。评估发现连续 VAE 的视觉表征优于离散 VT,后者常导致重建失真、细粒度纹理丢失以及文本和物体完整性受损。团队还对 GPT-4o 图像生成进行实验并讨论其潜在的 AR 特性,基准与代码已公开。
正文
Abstract:Autoregressive (AR) models have recently shown strong performance in image generation, where a critical component is the visual tokenizer (VT) that maps continuous pixel inputs to discrete token sequences. The quality of the VT largely defines the upper bound of AR model performance. However, current discrete VTs fall significantly behind continuous variational autoencoders (VAEs), leading to degraded image reconstructions and poor preservation of details and text. Existing benchmarks focus on end-to-end generation quality, without isolating VT performance. To address this gap, we introduce VTBench, a comprehensive benchmark that systematically evaluates VTs across three core tasks: Image Reconstruction, Detail Preservation, and Text Preservation, and covers a diverse range of evaluation scenarios. We systematically assess state-of-the-art VTs using a set of metrics to evaluate the quality of reconstructed images. Our findings reveal that continuous VAEs produce superior visual representations compared to discrete VTs, particularly in retaining spatial structure and semantic detail. In contrast, the degraded representations produced by discrete VTs often lead to distorted reconstructions, loss of fine-grained textures, and failures in preserving text and object integrity. Furthermore, we conduct experiments on GPT-4o image generation and discuss its potential AR nature, offering new insights into the role of visual tokenization. We release our benchmark and codebase publicly to support further research and call on the community to develop strong, general-purpose open-source VTs.
| Comments: | Accepted to AACL-IJCNLP 2026 (Main Conference). 26 pages, 13 figures, 8 tables. Code: this https URL. Dataset: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2505.13439 [cs.CV] |
| (or arXiv:2505.13439v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2505.13439 arXiv-issued DOI via DataCite |
Submission history
From: Huawei Lin [view email]
[v1]
Mon, 19 May 2025 17:59:01 UTC (21,050 KB)
[v2]
Thu, 1 Oct 2026 18:28:58 UTC (21,336 KB)
来源:arXiv:cs.LG · arxiv.org