arXiv:cs.AI· Saisab Sadhu, Shreeyans Arora, Pratinav Seth·· 3 小时前
法律思维链忠实性反事实审计:模型援引法条并非真正依据
Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
AI 导读
一项反事实审计发现,7 个开源权重模型(8B-70B)在 4 个法律推理基准上被要求援引裁判依据时,66.7%-100% 能给出正确法条,但替换法条后判决随之改变的比率仅 0.0%-21.7%(CaseHOLD)、30.0%-76.7%(ECHR 和 SCOTUS)、43.3%-50.0%(ContractNLI)。
正文
Abstract:Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY) |
| Cite as: | arXiv:2610.12361 [cs.AI] |
| (or arXiv:2610.12361v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12361 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Saisab Sadhu [view email]
[v1]
Thu, 8 Oct 2026 17:26:17 UTC (4,048 KB)
来源:arXiv:cs.AI · arxiv.org