arXiv:cs.AI· Rahul Sharma, Andrew B. Ducan, Ga\'etan Marceau Caron, Sebastian J. Vollmer·· 4 小时前
TypedBench:面向 System One 决策模型的校准、表述敏感性与成本基准
TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
AI 导读
TypedBench 是面向 Jev 等类型化决策模型的基准,基于 7 个策略标注生成器和 9 个评测套件构建,同时考察策略遵循、表述鲁棒性、概率质量与决策成本。评测覆盖一个托管模型、一个开源编码器及 0.8B-9B 参数的开源解码器:托管模型遵循策略但表述敏感且系统性欠自信,在非对称成本下用其概率反而不如直接取最高答案。解码器路由精确,但随选项或问题增多而变慢,在策略问题上准确率最低。
正文
Abstract:System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface. Software can act on these probabilities through thresholds, cost-weighted choices, and escalation rules. Consequently, if these probabilities are miscalibrated or wording-sensitive, the software ma take unintended actions leaving human operators with no textual rationale to inspect. Current evaluations largely report accuracy and calibration on public classification datasets without a clear reference. We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites. We report accuracy as median and range across paraphrases, and calibration error relative to the finite-sample noise floor of a matched, perfectly calibrated predictor. We assess probability quality through selective prediction, ordinal proper scoring rules, and realised cost under asymmetric cost matrices. We evaluate a hosted model, an open encoder, and a family of open decoders spanning 0.8B-9B parameters on identical items. The hosted model follows the stated policy but is wording-sensitive and systematically underconfident; under asymmetric costs, using its probabilities can be worse than taking its top answer. Decoders route exactly and are slow as options or questions are added. The decoder is least accurate on policy questions and degrades with more options. Overall, typed decision models must be evaluated jointly on policy adherence, wording robustness, probability quality, and induced decision outcomes.
| Subjects: | Artificial Intelligence (cs.AI); Applications (stat.AP) |
| Cite as: | arXiv:2610.11392 [cs.AI] |
| (or arXiv:2610.11392v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11392 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rahul Sharma [view email]
[v1]
Thu, 8 Oct 2026 07:22:36 UTC (219 KB)
来源:arXiv:cs.AI · arxiv.org