跳到正文
arXiv:cs.LG· Xiaocong Yang·· 2 天前AI 评分28

从事后解释到生成式可解释性:神经符号模型如何让推理过程可审计

Generative Interpretability via Scalable Neuro-Symbolic Models

AI 导读

针对 LLM 从聊天机器人走向智能体后事后可解释性无法在推理输出前审计或干预的问题,研究者提出"生成式可解释性"这一架构属性,使模型推理过程原生暴露语义可理解的检查点并支持因果干预,并以神经符号模型作为具体实现方案。该工作发表于 ACM AI Summit 2026,arXiv 编号 2609.13529。

正文

View PDF HTML (experimental)

Abstract:As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}, an architectural property under which a model's inference pass natively exposes semantically meaningful checkpoints that are human-understandable and amenable to causal intervention. We show the merits of generative interpretability as comparison to other interpretability research paradigms, and propose Neuro-Symbolic Models as a concrete instantiation.
Comments: ACM AI Summit 2026
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Symbolic Computation (cs.SC)
Cite as: arXiv:2609.13529 [cs.LG]
  (or arXiv:2609.13529v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.13529

arXiv-issued DOI via DataCite

Related DOI: https://doi.org/10.1145/3806096.3844850

DOI(s) linking to related resources

Submission history

From: Xiaocong Yang [view email]
[v1] Fri, 11 Sep 2026 20:54:08 UTC (43 KB)
[v2] Thu, 1 Oct 2026 04:08:19 UTC (284 KB)

来源:arXiv:cs.LG · arxiv.org