跳到正文
arXiv:cs.LG· Amir Rafe, Subasish Das·· 4 小时前AI 评分34

Kumo Tabular 与 Jev 如何估计德州车祸记录中未编码的碰撞因素

Estimating Uncoded Crash Factors with Tabular Foundation and System One Models: Kumo Tabular and Jev

AI 导读

该研究用表格基础模型 Kumo Tabular 读取 2017 至 2025 年 5,601,890 起德州车祸的编码记录,并用校准的 System One 模型 Jev 阅读两个概率样本的叙述文本,再以人工判断重校准概率,通过多轮 predict-then-debias 估计器合并三层结果。

正文

View PDF HTML (experimental)

Abstract:Road safety programs count the coded fields of police crash records, while the officer's narrative, which often records factors the fields omit, is rarely read. A safety office thus cannot tell how much its counts miss or where to review. This study develops and evaluates a system that joins both views of the 5,601,890 Texas crashes from 2017 to 2025 into population estimates with stated validity. An in-context tabular foundation model, Kumo Tabular, reads the coded record of every crash, a calibrated System One model, Jev, reads the narratives of two probability samples, and human judgments recalibrate its probabilities. A multiwave predict-then-debias estimator joins the three tiers, and a second human tier drawn with recorded probabilities checks the estimates by design. For hydroplaning, medical episodes, fatigue, animals, and phone use, the narrative documents more injury crashes than the coded field, 15,074 against 7,340 for phone use, and the human check agrees with all fifteen estimates within its margin. A re-read list ranked by Kumo Tabular finds confirmed discordance 7 to 58 times as often as random reading. At the planning cost of human coding, one further round of human judgments would cut the root mean square relative half-width from 22.0 to 16.2 percent, against 21.2 for reading every narrative. Two calibrated readers of different views, joined by a sampling design, give a safety office counts, a discordance map, a validated re-read list, and a reading budget, with Kumo Tabular reading the table at 15 times the speed of TabPFN 3.5.
Comments: 26 pages, 9 figures, 8 tables. Code: this https URL
Subjects: Applications (stat.AP); Machine Learning (cs.LG); Methodology (stat.ME)
MSC classes: 62D05, 62P30
ACM classes: I.2.6; G.3
Cite as: arXiv:2610.10321 [stat.AP]
  (or arXiv:2610.10321v1 [stat.AP] for this version)
  https://doi.org/10.48550/arXiv.2610.10321

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Amir Rafe [view email]
[v1] Wed, 7 Oct 2026 16:12:08 UTC (1,863 KB)

来源:arXiv:cs.LG · arxiv.org