arXiv:cs.LG· Arkanath Pathak, Unnat Jain, Alexander C. Berg·· 5 小时前AI 评分47
LLMAE:将预训练 LLM 改造为高保真连续文本自编码器
Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders
AI 导读
LLMAE 通过将预训练 decoder-only LLM 中间层激活暴露为定长潜空间瓶颈,配合结构化注意力掩码与 LoRA 适配,把 LLM 改造成连续文本自编码器。
正文
Abstract:Next-token prediction has enabled highly fluent autoregressive language models, but it represents global structure only indirectly through sequential factorization. In contrast, high-fidelity autoencoders have become a standard primitive in image generation, enabling generative models to operate over continuous latent spaces; text lacks a comparably faithful continuous representation. We propose LLMAE, a method for repurposing a pretrained decoder-only language model as a continuous text autoencoder by exposing the activations of an intermediate layer as a fixed-length latent bottleneck. LLMAE achieves this interface with structured attention masks and LoRA adaptation, leveraging the generative prior of the original LLM. Across two backbones (270M Gemma 3 and 0.5B Qwen2.5), LLMAE reconstructs sequences up to 1024 tokens with high accuracy, reproducing up to 97% of documents verbatim. We demonstrate downstream utility by training a lightweight diffusion model that generates detailed image captions directly in the frozen LLMAE latent space. By mapping text into a fixed-length continuous latent space, our approach provides an effective substrate for downstream adaptation while benefiting from the fluency of the original LLM. Code and models are available at this https URL.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.27248 [cs.LG] |
| (or arXiv:2609.27248v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.27248 arXiv-issued DOI via DataCite |
Submission history
From: Arkanath Pathak [view email]
[v1]
Wed, 23 Sep 2026 02:32:12 UTC (1,616 KB)
[v2]
Fri, 2 Oct 2026 04:25:28 UTC (1,624 KB)
来源:arXiv:cs.LG · arxiv.org