arXiv:cs.AI(全量分类)· Zhengkai Tu, Mingda Zhang, Zijia Wang, Xiaoying Tang, Jimmy Huang·· 5 小时前AI 评分41
JusticeAxis:法律判决在刚性规则与无依据裁量之间的基准测试
JusticeAxis: Benchmarking Legal Judgment between Rigid Rule Application and Ungrounded Discretion
AI 导读
研究团队提出 JusticeAxis,包含来自 18 个国家的 256 个真实刑事案件,每案配有音频、图像和文本证据及三份律师撰写的判决书(记录判决与两种失败判决),将法律判决形式化为参考锚定任务。
正文
Abstract:A sound judgment applies the law to established facts and weighs the circumstances in which they arose. However, existing methods swing between rigid statute matching and ungrounded discretion, benchmarks score a label or a rubric, and the experience that would supply the balance stays unverified. We formalize legal judgment as a reference-anchored task, whose object is a single decision that stays tied to the statute and to the circumstances at once. We introduce JusticeAxis, 256 real-world criminal cases from 18 countries with audio, image, and text evidence, and three lawyer-written judgments for every case: the recorded one and one for each failure. We further propose JusticeAgent, a harness whose element agents establish the facts and whose judge agent applies the law under skills carrying experience of the circumstances. Skills are distilled from execution trajectories and admitted only under Bayesian credible bounds. Experiments show that failure turns direction with scale: open-weight backbones drift to unsupported grounds, frontier models to the statutory default. We further verify that JusticeAgent, as a simple yet effective plugin, carries a frozen open-weight backbone to commercial level. Project resources are available at this https URL.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00353 [cs.AI] |
| (or arXiv:2610.00353v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00353 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mingda Zhang [view email]
[v1]
Wed, 30 Sep 2026 00:17:18 UTC (377 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org