arXiv:cs.CL· Yuyang Zhao, Xuan Liu, HaoYang Shang, Haojian Jin·· 5 小时前AI 评分36
测试时缩放与训练在个体立场预测中为何失效?
Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
AI 导读
研究在 STANCE-BENCH(2499 个来自 500 名 Hacker News 用户的预测任务)上评估测试时缩放与后训练方法,识别出错误共识、选择失败、响应过拟合和早期平台四种失效模式。结合候选立场直接打分与个人历史支持度评估的方法,在 781 任务测试集上用 Qwen3-8B 取得 21.83 的 Macro F1,高于直接打分的 19.27。
正文
Abstract:Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms prediction; and (4) early plateau, where reinforcement learning shows modest initial gains followed by limited further improvement. We expose these failures using STANCE-BENCH, which contains 2499 prediction tasks from 500 Hacker News users. Guided by this analysis, we explore a simple approach that combines direct scores for all candidate stances with explicit assessments of support from the individual's history. On the 781-task test set, this approach achieves 21.83 discussion-specific Macro F1 with Qwen3-8B, compared with 19.27 for direct scoring. Our results motivate evaluating candidate generation, final selection, and person-specific evidence use separately. Our data is available at this https URL.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.33155 [cs.CL] |
| (or arXiv:2609.33155v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33155 arXiv-issued DOI via DataCite |
Submission history
From: HaoYang Shang [view email]
[v1]
Sun, 27 Sep 2026 03:27:05 UTC (980 KB)
[v2]
Tue, 6 Oct 2026 07:37:15 UTC (1,021 KB)
来源:arXiv:cs.CL · arxiv.org