arXiv:cs.AI· Pratyay Dutta, Kowshik Thopalli, Vivek Narayanaswamy·· 6 小时前AI 评分41
COMPASS:在语言模型中定位推理能力所在
COMPASS: Finding Where Reasoning Lives in Language Models
AI 导读
COMPASS 是一种推理时激活引导方法,通过 logit 空间归因分数识别可干预的注意力头,并沿“答案正确性”方向引导其激活。在三个模型家族和多个数学基准上,它相较激活引导基线表现更优,GSM8K 准确率平均提升 16 个百分点,并以少 20-70% 的生成 token 接近 CoT 准确率。干预无需重新拟合即可迁移到未见基准,消融实验表明正确性方向与少数关键注意力头缺一不可。
正文
Abstract:Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoning features. For mathematical reasoning with verifiable answers, we show that a much simpler signal suffices, which is the correctness of the model's own direct answer attempts. This signal yields a latent direction that elicits reasoning. This direction is decodable within the activations of most attention heads, but only a small subset of them can be effectively intervened. We introduce COMPASS, an inference-time steering method that identifies these heads using a logit-space attribution score and steers their activations along the correctness direction, requiring only per-head activation statistics. Across three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines we compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70\% fewer generated tokens. Interventions transfer without re-fitting to unseen benchmarks, and ablations show that both the correctness direction and the small set of heads carrying it are necessary, with the effect concentrated in remarkably few heads.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07469 [cs.AI] |
| (or arXiv:2610.07469v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07469 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pratyay Dutta [view email]
[v1]
Mon, 5 Oct 2026 22:35:47 UTC (213 KB)
来源:arXiv:cs.AI · arxiv.org