跳到正文
arXiv:cs.LG· Mengzhe Geng·· 3 小时前

三通道评分冲突中的受控获取与弃权策略研究

Controlled Acquisition and Abstention in Three-Channel Score Conflicts

AI 导读

一项三评分基准研究提出阈值策略,在观察两个评分后可付费请求第三个或选择弃权。在部分留出的合成数据上,该策略取得 0.789±0.006 的目标决策准确率和 0.481±0.014 的效用,优于三评分多数参考的 0.626±0.008 和 0.252±0.016。但匹配预算测试显示,选择器的效用增益依赖预算与已观察评分对的排序,在 50% 和 63.7% 预算下反而降低效用。

正文

View PDF HTML (experimental)

Abstract:When audio, video, and text disagree, accuracy alone does not show whether to acquire another source or abstain. We study these choices in a controlled three-score benchmark: a policy observes two signed scores, may request the third at a cost, and can abstain. The primary reward is mechanism-specific: abstention is correct only for one designated ambiguity mechanism and is penalized under mixed corruption. Matched controls show that a threshold policy matches always-request decisions with fewer requests; its advantage over always-answer fusion depends on the reward assigned to that ambiguity. On a partially held-out synthetic split, the threshold policy reaches 0.789 +/- 0.006 targeted decision accuracy and 0.481 +/- 0.014 utility across 83 seeds. A three-score majority reference reaches 0.626 +/- 0.008 and 0.252 +/- 0.016, but uses more information. In a matched-budget test, a train-only value selector improves utility over no-query and matched-random policies at 10% and 25% budgets, while pair uncertainty has higher utility at every budget. At 50% and 63.7% budgets, the selector lowers utility despite slightly higher non-ambiguous accuracy. If all abstentions are scored incorrect, majority outranks the threshold policy in utility. At a central temporal setting, full-trace controls match the neural models while position perturbations separate them. On held-out-actor emotion clips, eight-frame fusion has opposite-signed accuracy differences for two encoder pairs, with both actor intervals containing zero; matched-request routing gains are small and uncertain. These results separate full-modality accuracy from pre-request selection value and show that selection value depends on budget and the observed-pair ranking.
Subjects: Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2610.10808 [cs.LG]
  (or arXiv:2610.10808v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10808

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mengzhe Geng [view email]
[v1] Wed, 7 Oct 2026 19:09:30 UTC (101 KB)

来源:arXiv:cs.LG · arxiv.org