arXiv:cs.CL· Linqi Zhang, Chong Qi, Yan Cheng, Wanqing Cao, Yu Liu, Chenwei Lin, Xian Xu·· 3 小时前AI 评分41
InsClaimBench:面向决策链的保险理赔裁决基准
InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain
AI 导读
研究者推出端到端基准 InsClaimBench,覆盖车险、财产险和健康险,含 375 个案件家族的 3,780 个案例与 86,656 条原子规则判断,逐层评估规则、裁决模块到赔付决策与金额。
正文
Abstract:Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.
| Comments: | 17 pages, 3 figures, 11 tables |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.09671 [cs.CL] |
| (or arXiv:2610.09671v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09671 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chenwei Lin [view email]
[v1]
Wed, 7 Oct 2026 08:35:23 UTC (449 KB)
来源:arXiv:cs.CL · arxiv.org