arXiv:cs.AI· Geonwoo Bang, Dongho Kim, Moohong Min·· 9 小时前AI 评分32
EVOL:模拟器引导的进化专家合成实现免部署学习路径推荐
EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation
AI 导读
EVOL 用知识追踪模拟器通过进化搜索合成每位学习者的专家示范,并训练一个免部署的前馈策略,在 ASSIST15、Junyi、EdNet 三个数据集和 L=5、10、20 路径长度上超越 8 个基线方法。
正文
Abstract:Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concepts without intermediate feedback, producing a combinatorial search space that grows super-exponentially with L and provides reward only at the final step. Second, expert learning paths would be the natural cure for sparse-reward RL, but they do not exist in educational data, because student logs record what learners did, not what they should have done. We address both obstacles by importing a recipe from simulator-based demonstration learning in robotics: the knowledge tracing simulator is used both to synthesize per-learner expert demonstrations through evolutionary search and to train a deployment-free policy that distills these demonstrations into a feed-forward learner. Our framework, EVOL, instantiates this pipeline with an asymmetric actor-critic where the actor commits to deployment-realistic blind planning while the critic exploits the privileged simulator state during training. Across three datasets (ASSIST15, Junyi, and EdNet; 39-189 concepts) and path lengths L = 5, 10, and 20, EVOL surpasses 8 baselines spanning heuristic, sequential, RL, graph-enhanced RL, and LLM-enhanced methods. We further compare three imitation strategies (BC, AWR, and DAPG) and show that final performance is governed by the quality of evolutionary experts rather than by the particular imitation objective.
| Comments: | Accepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026) |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03273 [cs.AI] |
| (or arXiv:2610.03273v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03273 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Geonwoo Bang [view email]
[v1]
Fri, 2 Oct 2026 13:16:40 UTC (220 KB)
来源:arXiv:cs.AI · arxiv.org