跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Barproda Halder, Qiuyi Zhang, Sanghamitra Dutta·· 13 小时前AI 评分38

用部分信息分解解读大语言模型的推理过程

Interpreting Reasoning of Large Language Models via Partial Information Decomposition

AI 导读

研究者提出可解释性框架 SLIDER,利用部分信息分解将相邻推理步骤中关于最终答案的信息拆分为唯一、冗余和协同三类非负成分,并据此定义 Step-RRI 指标。

正文

View PDF HTML (experimental)

Abstract:Large reasoning models (LRMs) have achieved substantial improvements in solving complex mathematical problems, but often produce lengthy, repetitive, or erroneous reasoning trajectories. In this work, we introduce a new interpretability framework, SLIDER, to evaluate the quality of the reasoning process. SLIDER leverages an emerging body of work from information theory called Partial Information Decomposition to disentangle the information about the final answer between two consecutive reasoning steps into non-negative components: unique information (in preceding steps or current step), redundant information, and synergistic information. Building on this decomposition, we propose the *Step-wise Repetitive Reasoning Index (Step-RRI)*, a theoretically grounded measure that assesses whether the answer-relevant information in the current step $S_i$ is predominantly redundant with the past steps $S_{<i}$, relative to its unique and synergistic contributions. To evaluate the effectiveness of Step-RRI in detecting repetitiveness, we apply SLIDER to the redundancy class of the PRMBench dataset where Step-RRI improves step-level redundancy identification accuracy by over $10$ points compared to embedding-similarity and information-gain baselines. Next, we define *Trajectory-RRI*, an aggregate measure of repetitiveness for an individual reasoning trajectory. To demonstrate its practical relevance, we show that average Trajectory-RRI strongly correlates with actual reasoning length across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and GPT-4.1, motivating its use as a signal for improving reasoning efficiency. Finally, we introduce *Trajectory-RRI-guided data selection for fine-tuning*, demonstrating that selecting training data based on Trajectory-RRI can improve a fine-tuned model's reasoning efficiency while largely preserving its task performance.
Comments: Accepted at ICLR 2026 Workshop on Logical Reasoning of Large Language Models
Subjects: Information Theory (cs.IT); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG)
Cite as: arXiv:2610.00571 [cs.IT]
  (or arXiv:2610.00571v1 [cs.IT] for this version)
  https://doi.org/10.48550/arXiv.2610.00571

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Barproda Halder [view email]
[v1] Wed, 30 Sep 2026 18:44:13 UTC (4,160 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org