跳到正文
arXiv:cs.AI· Jiawen Lu, Tongtong Wu·· 5 小时前AI 评分36

Laya 与 Jev 在类型化决策模型中的候选覆盖基准测试

Benchmarking Candidate Coverage in Typed Decision Models

AI 导读

研究提出配对候选覆盖基准协议,在 AG News、DBpedia、Emotion 和 TREC 上评估 Laya 与 Jev,每个模型基于 300 条校准文本和 589 条测试文本产生 23,932 次预测。

正文

View PDF HTML (experimental)

Abstract:Typed decision models return choices or distributions over answer options supplied at request time. Accuracy with complete options does not establish whether a model recognizes that a reference answer is missing or avoids rejecting valid candidates. We present a paired candidate-coverage benchmark protocol and an initial evaluation of Laya and Jev across AG News, DBpedia, Emotion, and TREC. The models receive identical frozen texts and requests: 300 calibration and 589 test texts yield 23,932 predictions per model. Present/absent pairs match ordinary candidate count, and name variants preserve descriptions, members, and order. Native rejection behavior differs sharply: at five TREC candidates with natural names, Laya detects 97.2% of missing-answer cases but falsely rejects 69.7% of present controls; Jev's rates are 24.8% and 0.0%. Calibration-only none-score thresholds change these rates to 33.9%/3.7% and 45.0%/1.8%, respectively. On DBpedia, Jev's high coverage-score AUROC supports a stronger operating point, whereas both models have weak complete-set accuracy on Emotion. Competence-conditioned analysis, probability-precision sensitivity, and interface audits show why classification, score ranking, and rejection policies need separate measurement. This initial benchmark is descriptive and limited to reference-label omission; it does not establish natural out-of-scope generalization, causal mechanisms, or a new rejection method.
Comments: 19 pages, 1 figure, 8 tables
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.03387 [cs.AI]
  (or arXiv:2610.03387v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03387

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jiawen Lu [view email]
[v1] Fri, 2 Oct 2026 14:36:46 UTC (64 KB)

来源:arXiv:cs.AI · arxiv.org