arXiv:cs.LG(机器学习,全量分类)· Amir Rafe, Subasish Das·· 14 小时前AI 评分39
系统一决策模型与训练分类器及语言模型在自动决策门中的基准对比
Benchmarking System One decision models against trained classifiers and language models for automated decision gates
AI 导读
一项基准测试在匹配条件下对比了六大家族八个决策模型检查点(含托管模型 Jev)与监督分类器、零样本分类器及生成式语言模型。有标签时小型训练分类器在意图任务上最准,无标签时除编码器类检查点外所有决策模型均超过零样本蕴含分类器。意图训练的第一阶段升级到 Jev 可在全 GPU 利用率下以 0.43 成本达到同等准确率。
正文
Abstract:Software that hands branching decisions to a model needs a declared option and a probability it can threshold. Typed decision models, also called System One models, return such probabilities without generating text, while supervised classifiers and generative language models are the established alternatives. Under matched conditions, one harness sends eight decision-model checkpoints from six families, including the hosted model Jev, and two generative comparators the same semantic requests, and scores supervised and zero-shot classifiers on the same workflow, intent and social-science items. The ranking of the model classes depends on the conditions. With the task's own labels, small trained classifiers are the most accurate on intents and not significantly different from the best decision models on workflows. Without labels, every decision model except the encoder-based checkpoints exceeds a zero-shot entailment classifier on workflows and intents. Read through option-key likelihoods, a larger generative model is level with Jev on workflows and intents and accepts more workflow decisions at five percent risk, and fine-tuned decision checkpoints gain intent accuracy over their untuned backbones. Stored temperatures fitted on few options raise calibration error with many options, and a held-out threshold for five percent in-scope risk still lets Jev accept 0.310 of out-of-scope requests. Swapping yes and no flips 50.5 answers per hundred for Jev, while fine-tuned checkpoints cut their backbones' social-science flips. An intent-trained first stage escalating to Jev matches its accuracy at 0.43 of its cost at full graphics-processor utilization. The results yield condition-dependent design rules for automated decision gates.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00346 [cs.LG] |
| (or arXiv:2610.00346v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00346 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Amir Rafe [view email]
[v1]
Tue, 29 Sep 2026 19:57:59 UTC (612 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org