arXiv:cs.LG· Debjyoti Saha Roy·· 4 小时前AI 评分46
Characterize Then Distill:大输出空间中的机制化推理研究
Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces
AI 导读
研究提出 MISTILL 蒸馏方法,在监督学生模型推理文本之外,还对决策事件处的池化注意力加以监督。在 MIMIC-IV 临床编码任务中,将全部 5,651 个候选诊断编码放入上下文时,一小组全局、分阶段结构的注意力头被消融与 knock-in 实验证明是必要且充分的。跨模型族学生模型的对比决策因果恢复率接近翻倍,在已具备大部分该能力的模型上也有小幅稳定增益。
正文
Abstract:Reasoning-trained language models can perform, zero-shot, multi-label tasks that require selecting a small set of relevant labels from a universe of thousands to hundreds of thousands of candidates. We ask how they do it mechanistically, and whether the mechanism can be distilled. We make the question measurable by treating each decision as a token-level event scored by the model's own decision margin: the token that picks a coarse region of the label space, the tokens that pick a label within it, and the token where the output departs from a close alternative (a near-miss) named earlier in the reasoning. Attribution, exact mean-ablation, knock-in into another example's context, and a null calibration that discounts generic heads then give individual attention heads causal standing. On clinical coding of hospital discharge summaries (MIMIC-IV), with all 5,651 candidate diagnosis codes in context, a small, global, phase-structured set of heads is necessary and sufficient, by ablation and knock-in, on essentially every summary; distinct head families attend to the candidate region and back to the near-miss named earlier; and, for the mentions decided in the reasoning, the region can already be elicited several tokens before the code, from a disjoint mid-layer set that reads the input. We introduce MISTILL: unlike chain-of-thought distillation, which transfers only the teacher's reasoning text, it also supervises the student's pooled attention at exactly these decision events. Read on heads found after training, it nearly doubles the causal recovery of the contrastive decision in a cross-family student and adds a small, seed-stable gain in one that already carries most of it, with no detected task difference when both objectives train bf16 weights and a task cost with fp32 master weights.
| Comments: | substantially revised and extended; supersedes v1. New analysis (token-level decision events, head-level causal tests on MIMIC-IV clinical coding), new distillation method (MISTILL), experiments and text; the author list reflects authorship of this version. 58 pages, 6 figures. Code: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.06840 [cs.CL] |
| (or arXiv:2606.06840v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.06840 arXiv-issued DOI via DataCite |
Submission history
From: Debjyoti Saha Roy [view email]
[v1]
Fri, 5 Jun 2026 02:32:24 UTC (334 KB)
[v2]
Tue, 6 Oct 2026 02:10:23 UTC (301 KB)
来源:arXiv:cs.LG · arxiv.org