arXiv:cs.LG· Tingzhu Bi, Ping Wang, Meng Ma·· 3 小时前AI 评分53
arXiv 论文 Nautil:教 LLM 调查员判断何时结案
Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case
AI 导读
论文研究 LLM 调查员在事故、缺陷与故障调查中判断证据是否足以结案的问题,发布 Nautil 数据集,含 731 个经审计的案例,覆盖航空、铁路、海事、化学品安全和车辆缺陷报告及生产服务器事故,附教师轨迹、分布外测试集和反事实证据版本。
正文
Abstract:Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is "cause undetermined". Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control, overstatement falls from 97% to 35%, and correct, non-overstated conclusions rise from 3% to 43%. Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence.
| Comments: | 23 pages. Dataset: this https URL ; Models: this https URL , this https URL ; Demo: this https URL ; Code: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03190 [cs.CL] |
| (or arXiv:2610.03190v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03190 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tingzhu Bi [view email]
[v1]
Fri, 2 Oct 2026 12:06:02 UTC (826 KB)
来源:arXiv:cs.LG · arxiv.org