arXiv:cs.AI(全量分类)· Sam Larson·· 5 小时前AI 评分33
对话幽默训练中的奖励漏洞与对策:Comedic Fool's Gold 研究
Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor
AI 导读
研究揭示对话幽默训练中两类自动奖励的漏洞:基于嵌入向量的"惊喜"奖励会把词序打乱的回复与机智回复同等接受,加入流畅度过滤器后虽能识别乱序,却也会误拒部分机智回复;观众模型预测的笑声则易受对话双方笑声线索影响,跨说话人归一化可封堵已覆盖攻击,但未匹配的表达仍可被利用。三轮强化学习实验中,最终一轮将综合评估分提升 0.0903、零分会话减少 40%,但幽默专项提升仍低于预注册目标。
正文
Abstract:We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency filter detects the shuffles, but the combined reward also rejects some witty replies and fails further validation. An audience model's predicted laughter is instead vulnerable to laughter cues in either speaker's messages. Normalizing these cues across speakers blocks the covered attacks, although unmatched expressions remain exploitable. Three reinforcement-learning runs evaluate training with successive reward revisions. The final run improves the combined evaluation score by 0.0903 and reduces zero-score sessions by 40%, but its humor-specific improvement remains below our preregistered target. These findings illustrate a broader challenge for automated reward design: countermeasures must block exploitable shortcuts while preserving the behavior the reward was intended to encourage.
| Comments: | 11 pages, 3 figures, 4 tables |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00197 [cs.AI] |
| (or arXiv:2610.00197v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00197 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Samuel Larson [view email]
[v1]
Fri, 18 Sep 2026 20:14:09 UTC (56 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org