arXiv:cs.AI· Hao Jiang, Xin Li, Annan Wang, Yichi Zhang, Weisi Lin·· 9 小时前AI 评分55
HACKTRACE:在代码生成中基于行为监督检测 reward hacking
hacktrace: behavior-supervised detection of reward hacking during code generation
AI 导读
论文发布 HACKTRACE,一种行为监督监测器,通过读取编码智能体生成代码时已计算的内部状态来检测 reward hacking,包括未成功的作弊尝试。
正文
Abstract:A coding agent can earn a passing grade by fixing its code, or by deleting the test that exposes the bug. Detecting such reward hacking requires recognizing attempted shortcuts, including those that fail. We release 173,561 annotated multi-turn coding trajectories from Qwen3-8B and show that supervising shortcut behavior independently of exploit success substantially improves detection. We introduce HACKTRACE, a behavior-supervised monitor that reads the internal states the agent already computes while generating code. Reusing these states enables monitoring before a turn is complete, without additional language-model tokens or passes. Combining this evidence with static features of the final files achieves a mean per-problem AUC of 0.997 with 8 ms of monitoring overhead, improving both accuracy and latency over monitors that run the model again on an honesty question and answer. The same generation states also provide an inexpensive monitoring signal for reinforcement learning. With strong GRPO penalties, HACKTRACE reduces the cheating share of passing solutions from 82-91% to 1-5%, while retaining honest, correct solutions and maintaining high detection accuracy as the policy evolves. Our results show that both the supervision target and the source of monitoring evidence matter for turning accurate detection into a useful training signal.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03055 [cs.AI] |
| (or arXiv:2610.03055v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03055 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hao Jiang [view email]
[v1]
Fri, 2 Oct 2026 09:35:59 UTC (433 KB)
来源:arXiv:cs.AI · arxiv.org