arXiv:cs.AI· Rub\'en Manrique, Michelle Castellanos, Jorge Morales, Juan David Guti\'errez, Antonio Barreto Rozo, Joaqu\'in V\'elez Navarro·· 5 小时前AI 评分45
LLM 懂哥伦比亚法律吗?面向哥伦比亚法律体系的可靠性基准测试
Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System
AI 导读
研究者推出经专家验证的哥伦比亚法律基准,含 1042 道题、覆盖 10 个法律领域和三种题型,评测 15 个专有与开源模型。闭卷选择题准确率从 0.905(Gemini 3.1 Pro)到 0.577,但自由文本法律答案的事实正确率均不超过 0.45,且答案相关性与正确性呈负相关(rho = -0.46)。模型引用的法条仅约一半正确,研究认为当前 LLM 处理哥伦比亚法律任务仍需专家监督。
正文
Abstract:Large language models (LLMs) are increasingly used to support legal practice, education, and research, yet their reliability in national legal systems outside the United States remains largely undocumented. We introduce an expert-validated benchmark for evaluating LLM reliability on the Colombian legal system. The benchmark comprises 1,042 items spanning ten areas of law and three question formats (closed multiple-choice, semi-open, and open-ended IRAC), built through a human-in-the-loop pipeline with multi-stage expert review. We evaluate 15 contemporary proprietary and open-weight models with format-appropriate metrics. Accuracy on closed questions ranges widely, from 0.905 (Gemini 3.1 Pro) to 0.577, but on free-text legal answers factual correctness never exceeds 0.45 (on a 0-1 scale) for any model. We find a dissociation between answer relevancy and correctness (Spearman rho = -0.46): models reliably sound responsive while frequently being wrong, a pattern of particular concern for non-expert users. Closed-question accuracy and free-text correctness are strongly rank-correlated (rho = 0.94), so cheap multiple-choice screening predicts model ranking but overstates absolute reliability. An independent rubric-based LLM judge and blind human expert scoring both reproduce the free-text ranking (rho >= 0.88). The judge further reveals that only about half of the norms models cite are correct; the rest are wrong or non-existent. Reliability varies systematically by legal area and follows an inverted-U across question complexity. Our results indicate that current LLMs require expert supervision for Colombian legal tasks, and that grounding answers in authoritative sources is a promising path to higher reliability. We release the benchmark construction pipeline to support reproducible evaluation.
| Comments: | 38 pages, 23 figures, 8 tables |
| Subjects: | Artificial Intelligence (cs.AI); Computers and Society (cs.CY) |
| ACM classes: | I.2.7; K.5.1 |
| Cite as: | arXiv:2610.03639 [cs.AI] |
| (or arXiv:2610.03639v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03639 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rubén Manrique [view email]
[v1]
Fri, 2 Oct 2026 17:27:41 UTC (292 KB)
来源:arXiv:cs.AI · arxiv.org