arXiv:cs.LG(机器学习,全量分类)· Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu·· 14 小时前AI 评分44
Faynt:面向《任天堂明星大乱斗 Melee》竞技对战的策略扩展与优化
Faynt: Scaling and Optimizing Policies for Competitive Melee
AI 导读
研究团队推出 Faynt,一组 10M 和 75M 参数的 Transformer 策略,单个 checkpoint 即可操控《Super Smash Bros. Melee》全部 26 个角色。经强化学习后,10M 模型在同角色对战中赢下 244 局中的 240 局(98.4%),并在对零延迟 Slippi-AI 的 68 局评测中全胜。团队开源了权重、两套基准测试及自动化锦标赛平台。
正文
Abstract:We introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee, each controlling all 26 characters with a single checkpoint. After reinforcement learning (RL), the 10M wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases on their supported rosters, with a winning record against every release. These opponents retain 21- or 24-frame action delays; Faynt uses no added delay, and we have not isolated the effect of this difference. In a separate evaluation against a privately supplied zero-delay Slippi-AI model, the 10M wins all 68 games across two conditioning settings. We study architecture, optimization, scaling, and hyperparameter transfer to guide pretraining on approximately 840,000 human replays. Post-training combines rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches. On the initial 152-game benchmark, the supervised 10M wins 69.7% of games, compared with 45.4% for the pretrained 75M, despite higher overall held-out controller-prediction loss. The weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies. After supervised post-training, both models take less damage per minute, build larger early leads, and win more often after losing the first life. Optimized inference on recorded game states averages 5.2 ms per decision for the 10M and 8.7 ms for the 75M on an NVIDIA T4, excluding emulator execution and communication. We open-source the weights, both benchmark suites, and a platform for automated model tournaments.
| Comments: | 54 pages. Preprint, in review |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02144 [cs.LG] |
| (or arXiv:2610.02144v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02144 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ali Janati [view email]
[v1]
Thu, 1 Oct 2026 17:48:38 UTC (858 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org