arXiv:cs.LG· Xiaocong Yang·· 2 天前AI 评分28
从事后解释到生成式可解释性:神经符号模型如何让推理过程可审计
Generative Interpretability via Scalable Neuro-Symbolic Models
AI 导读
针对 LLM 从聊天机器人走向智能体后事后可解释性无法在推理输出前审计或干预的问题,研究者提出"生成式可解释性"这一架构属性,使模型推理过程原生暴露语义可理解的检查点并支持因果干预,并以神经符号模型作为具体实现方案。该工作发表于 ACM AI Summit 2026,arXiv 编号 2609.13529。
正文
Abstract:As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}, an architectural property under which a model's inference pass natively exposes semantically meaningful checkpoints that are human-understandable and amenable to causal intervention. We show the merits of generative interpretability as comparison to other interpretability research paradigms, and propose Neuro-Symbolic Models as a concrete instantiation.
| Comments: | ACM AI Summit 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Symbolic Computation (cs.SC) |
| Cite as: | arXiv:2609.13529 [cs.LG] |
| (or arXiv:2609.13529v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.13529 arXiv-issued DOI via DataCite |
|
| Related DOI: | https://doi.org/10.1145/3806096.3844850
DOI(s) linking to related resources |
Submission history
From: Xiaocong Yang [view email]
[v1]
Fri, 11 Sep 2026 20:54:08 UTC (43 KB)
[v2]
Thu, 1 Oct 2026 04:08:19 UTC (284 KB)
来源:arXiv:cs.LG · arxiv.org