跳到正文
arXiv:cs.AI· Yi Ma, Tianpei Yang, Yaodong Yang, Weixun Wang, Hongyao Tang·· 3 小时前

Jev 能否作为强化学习中的 Q 函数或策略?

Can Jev be Your Q or Policy in Reinforcement Learning?

AI 导读

研究首次将冻结的决策模型 Jev 用于强化学习训练过程,把它放在参考策略、探索评判和回放评分三个位置。在九项 MiniGrid 任务和三个 Atari 游戏上,使用 Jev 训练的表现超过标准 RL 学习器,包括学习器单独训练毫无进展的场景,而 Jev 本身保持未训练状态。

正文

View PDF HTML (experimental)

Abstract:Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly. Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass. Existing work studies foundation models in RL either as models to be trained or as generators to be prompted, and Jev belongs to neither category, having so far served only as a black box in single domains. How well such a model decides on its own in RL environments, and how it can improve RL as a component of training, therefore remain unaddressed. To this end, in this paper we first examine the requirements that the objects of an RL system place on the answers they consume, and establish that Jev can fulfill all of them except the cardinal use of a value function. The remaining objects form positions that admit several roles each. We then construct algorithms that employ Jev at three of these positions, as a reference policy, an exploration judge, and a replay rater, to improve sample efficiency, exploration, and learning performance. Across nine MiniGrid tasks and three Atari games, training with Jev outperforms a standard RL learner, including where the learner makes no progress alone, while the model itself remains untrained. To our knowledge, we present the first use of Jev within the RL learning process and establish a frozen decision model as a usable component of RL training, inviting further exploration of how Jev and other advanced decision models can improve RL.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11692 [cs.LG]
  (or arXiv:2610.11692v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11692

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tianpei Yang [view email]
[v1] Thu, 8 Oct 2026 11:02:14 UTC (1,281 KB)

来源:arXiv:cs.AI · arxiv.org