arXiv:cs.CL· Hanhan Zhou, Shamik Roy, Rashmi Gangadharaiah·· 4 小时前AI 评分37
离散扩散语言模型的机制引导干预:自适应调度实现精准控制
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
AI 导读
研究者对四种离散扩散语言模型(124M-8B 参数)训练稀疏自编码器,发现不同属性在去噪过程中按不同时间表"定型"——如在 MDLM 上主题在前 2% 去噪步骤内即确定,而情感则需约 20% 的过程逐步浮现。基于此提出自适应调度机制,将干预集中在各属性活跃形成阶段。在四种 DLMs 的七项单属性与多属性控制任务中,该方法在控制力与生成质量权衡上均优于均匀干预和区间限制基线。
正文
Abstract:Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive models. Controlled generation methods for DLMs, imported from autoregressive models, apply uniform intervention at every denoising step. We show this uniform schedule is inefficient and degrades quality, and the damage compounds when multiple attributes are steered jointly. To diagnose the failure, we train sparse autoencoders on four DLMs (124M-8B parameters) and find that different attributes commit on distinct schedules, varying in timing, sharpness, and magnitude. For instance, topic commits within the first 2% of denoising on MDLM, whereas sentiment emerges gradually over 20% of the process. Motivated by these profiles, we propose an adaptive scheduling mechanism that concentrates intervention where each attribute is actively forming. An idealized allocation analysis predicts that attributes with more sharply concentrated emergence benefit more from adaptive scheduling, a prediction we confirm empirically. Across seven single- and multi-attribute steering tasks on four DLMs, adaptive steering consistently improves the control-quality balance over uniform and interval-restricted baselines.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2605.10971 [cs.LG] |
| (or arXiv:2605.10971v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.10971 arXiv-issued DOI via DataCite |
Submission history
From: Hanhan Zhou [view email]
[v1]
Fri, 8 May 2026 18:52:17 UTC (4,221 KB)
[v2]
Tue, 6 Oct 2026 20:54:30 UTC (4,337 KB)
来源:arXiv:cs.CL · arxiv.org