跳到正文
arXiv:cs.AI· Arman Behnam, Binghui Wang·· 9 小时前AI 评分52

CERTID 论文:推理模型在因果可识别性判断上准确但不可靠

Reasoning Models Are Accurate but Unsound on Identification

AI 导读

Arman Behnam 与 Binghui Wang 发布 CERTID,一个基于 ID 算法的因果可识别性认证与公式验证流程,并用结构因果模型精确验证返回公式。

正文

View PDF HTML (experimental)

Abstract:A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both limitations. CERTID uses the sound and complete causal identification algorithm ID to certify whether an effect is identifiable from a given graph and query, and verifies returned formulas against structural causal models whose interventional distributions are known exactly. CERTID further develops theoretical results to mitigate structural leakage, repair non-identifiable queries, and establish grading guarantees. We evaluate three frontier reasoning models (Gemini Flash, Gemini Pro, and GPT5.5) on 1,200 certified instances spanning 4 to 50 vertices. Accuracy proves a poor proxy for soundness: on identical instances, the false-claim rate on non-identifiable queries varies by seventeen-fold across models. We also find that models decide identifiability with 97-100% accuracy on graphs generated after the strongest model's training snapshot. Instances, the certification procedure, the verifier, and per-instance records are available at this https URL.
Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
Cite as: arXiv:2610.03519 [cs.AI]
  (or arXiv:2610.03519v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03519

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Binghui Wang [view email]
[v1] Fri, 2 Oct 2026 16:12:21 UTC (87 KB)

来源:arXiv:cs.AI · arxiv.org