跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Darpan Aswal, C\'eline Hudelot·· 7 小时前AI 评分33

DLIG:面向扩散语言模型的时序化 Token 归因方法

Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models

AI 导读

研究者提出 Diffusion Layer Integrated Gradients(DLIG),将 Integrated Gradients 扩展到扩散语言模型的任意层与去噪步,用于归因模型对输入提示词的渐进式承诺。

正文

View PDF HTML (experimental)

Abstract:This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.01177 [cs.CL]
  (or arXiv:2610.01177v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.01177

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Darpan Aswal [view email]
[v1] Thu, 1 Oct 2026 06:51:42 UTC (1,963 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org