Meta Superintelligence Labs 训练出 9B 用户模拟器 MIMESIS,基于真实人类对话与 13 种真实用户行为模式,行为保真度比 Claude Opus 5 高 13.4 分。
Banger paper from Meta Superintelligence Labs on user simulators for agent training.
Agent RL setups usually let an assistant LLM play the user, so the simulated user is too cooperative and too explicit.
A fixed GPT-5.5 agent finds tau-bench tasks easier with these users than with real people.
This work trains MIMESIS, a 9B user simulator, on human conversations and 13 behavior patterns observed in real users. It beats Claude Opus 5 on behavioral fidelity by 13.4 points.
Agents trained against it outperform agents trained against GPT-5.5 under all nine user simulators they never saw.
They find that adding a coaching step that turns the simulator's private reasoning into feedback adds further gains.
来源:DAIR.AI · x.com