arXiv:cs.LG· Harry Mayne, Justin Singh Kang, Dewi Gould, Kannan Ramchandran, Adam Mahdi, Noah Y. Siegel·· 6 小时前AI 评分53
论文提出 NSG 指标:LLM 自解释有助于预测模型行为
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
AI 导读
论文提出 Normalized Simulatability Gain(NSG)指标,以解释能否帮助观察者预测模型行为来衡量其忠实性。作者评估 Gemini 3、GPT-5.2、Claude 4.5 等 18 个前沿模型在健康、商业、伦理数据集的 7,000 个反事实样本,发现自解释将行为预测提升 11-37% NSG,且比外部模型生成的解释更具预测信息;同时 5-15% 的自解释高度误导。
正文
Abstract:LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, typically relying on identifying unfaithfulness via adversarial prompting or detecting reasoning errors. These methods overlook the predictive value of explanations. We introduce Normalized Simulatability Gain (NSG), a general and scalable metric based on the idea that a faithful explanation should allow an observer to learn a model's decision-making criteria, and thus better predict its behavior on related inputs. We evaluate 18 frontier proprietary and open-weight models, e.g., Gemini 3, GPT-5.2, and Claude 4.5, on 7,000 counterfactuals from popular datasets covering health, business, and ethics. We find self-explanations substantially improve prediction of model behavior (11-37% NSG). Self-explanations also provide more predictive information than explanations generated by external models, even when those models are stronger. This implies an advantage from self-knowledge that external explanation methods cannot replicate. Our approach also reveals that, across models, 5-15% of self-explanations are highly misleading. Despite their imperfections, we show a positive case for self-explanations: they encode information that helps predict model behavior.
| Comments: | ICML 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2602.02639 [cs.AI] |
| (or arXiv:2602.02639v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2602.02639 arXiv-issued DOI via DataCite |
Submission history
From: Harry Mayne [view email]
[v1]
Mon, 2 Feb 2026 18:54:51 UTC (1,151 KB)
[v2]
Tue, 6 Oct 2026 17:02:24 UTC (1,319 KB)
来源:arXiv:cs.LG · arxiv.org