跳到正文
arXiv:cs.AI· Saisab Sadhu, Shreeyans Arora, Pratinav Seth·· 3 小时前

法律思维链忠实性反事实审计:模型援引法条并非真正依据

Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness

AI 导读

一项反事实审计发现,7 个开源权重模型(8B-70B)在 4 个法律推理基准上被要求援引裁判依据时,66.7%-100% 能给出正确法条,但替换法条后判决随之改变的比率仅 0.0%-21.7%(CaseHOLD)、30.0%-76.7%(ECHR 和 SCOTUS)、43.3%-50.0%(ContractNLI)。

正文

View PDF HTML (experimental)

Abstract:Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
Cite as: arXiv:2610.12361 [cs.AI]
  (or arXiv:2610.12361v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12361

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Saisab Sadhu [view email]
[v1] Thu, 8 Oct 2026 17:26:17 UTC (4,048 KB)

来源:arXiv:cs.AI · arxiv.org