扩展并蒸馏文本嵌入以提升扩散模型可生成性
Scaling and Distilling Text Embeddings for Better Diffusibility
研究发现扩展文本嵌入能大幅提升连续扩散语言模型(DLM)性能,用 T5Gemma-2-270M 替换 ELF 模型中的 T5-small 嵌入,可将 Gen. PPL 降低约 40%。
Glad to share our recent paper on the latent space for continuous diffusion language models (DLMs)!
In this paper, we find that scaling text embeddings can greatly boost the performance of continuous DLMs; for instance, by replacing the T5-small embeddings used in the recent ELF models with the advanced T5Gemma-2-270M, we can reduce Gen. PPL by about 40%.
But scaling alone is not enough, as the scaled embeddings can be hard to generate. They are so distinctive and informative that embeddings of similar and interchangeable words are far apart. Thus, continuous diffusion struggles to generate such separated and discrete targets. To mitigate this, we distill them into a more connected and robust latent space, making them easier for diffusion to generate.
As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL. Our results shed light on language representation learning, and also on a reliable way to scale DLMs.
Code and project page are released too: code | project page. Welcome to check them out!
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co