跳到正文
arXiv:cs.AI· Kaisen Yang, Qingle Liu, Kejin Wang, Yicheng Zhao, Jieming Li, Shenghan Zheng, Ruize Yang, Bojun Yang, Heng Gong, Xiang Gao, Lanyue Zhang, Kaiyu Zhong, Zhuo Liu, Shaoxuan Li, Chengxi Li, Yong Yan, Weixuan Zhang, Tianwei Luo, Situ Wang, Youjie Zheng, Sihan Zhao, Shengyuan Wang, Huan-ang Gao, Jiazheng Xu, Xiaohui Xie, Wentao Han, Hongning Wang·· 3 小时前

AI 智能体能否通过启发式学习登顶?长时程游戏智能体竞赛中的启发式学习评估

Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition

AI 导读

研究者提出对抗式启发式学习(AHL),让 AI 智能体在不改变模型权重的前提下改写游戏策略与配套软件,并发布含 12 款对抗游戏、1920 份人类程序存档的 AAArena 基准。在评估中,Opus5.5 搭配 Claude Code 拿下 6 枚金牌,其余 6 个游戏人类天梯无人登顶,规则越复杂表现越弱。实验显示对手选择与密集反馈有助于策略改进。

正文

Authors:Kaisen Yang, Qingle Liu, Kejin Wang, Yicheng Zhao, Jieming Li, Shenghan Zheng, Ruize Yang, Bojun Yang, Heng Gong, Xiang Gao, Lanyue Zhang, Kaiyu Zhong, Zhuo Liu, Shaoxuan Li, Chengxi Li, Yong Yan, Weixuan Zhang, Tianwei Luo, Situ Wang, Youjie Zheng, Sihan Zhao, Shengyuan Wang, Huan-ang Gao, Jiazheng Xu, Xiaohui Xie, Wentao Han, Hongning Wang

View PDF HTML (experimental)

Abstract:Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.12341 [cs.AI]
  (or arXiv:2610.12341v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12341

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kaisen Yang [view email]
[v1] Thu, 8 Oct 2026 17:15:48 UTC (3,331 KB)

来源:arXiv:cs.AI · arxiv.org