arXiv:cs.AI(全量分类)· Cris Huynh·· 5 小时前AI 评分34
鲁棒即显著:知情对手将最优信号推向显著性极点
Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole
AI 导读
在知情对手共享受众的受限信号通道中,抗对手最优信号与既有工作中的显著性极点完全重合:108 个验证项上二者一致,20 万项池中仅 2748 项不同,且差异恰好出现在显著性-贝叶斯坐标未定义处。
正文
Abstract:When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool, the two differ on only 2,748 items --- lying exactly where the prior salience-to-Bayes coordinate is undefined. Where defined, robustness is achieved by moving from Bayesian discrimination entirely to salience. We show this by introducing an adversary to a forced-choice task (abstracted from Deception: Murder in Hong Kong). The adversary knows the target, observes the signal, and argues for the strongest wrong answer using a persuasion budget, $\beta$. As $\beta$ grows, the optimal signal shifts from the posterior-maximizing option to the margin-maximizing one; at $\beta = 0$, the game reproduces the original oracle model with a listener temperature of $\tau = 1$. This effect is real: 18.2 percent of the pool has an optimum that shifts under a finite budget, and each item's critical budget is exact. This coincidence structurally limits empirical evaluation. Two adversary framings change the chosen option of seven language models on 30 to 77 of 108 items against an exact no-effect rate. Yet, no measurement can determine whether this movement is toward the adversary-aware optimum or toward salience, because the two options are identical. This is a structural limit, not a null result. The diagnostic check is cheap: before evaluating adversary-awareness, verify whether the robust target coincides with a heuristic target on the evaluation items.
| Comments: | 11 pages, 3 figures |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00233 [cs.AI] |
| (or arXiv:2610.00233v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00233 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Cris Huynh Mr [view email]
[v1]
Wed, 23 Sep 2026 00:52:38 UTC (216 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org