arXiv:cs.CL· Harshita Chopra, Kshitish Ghate, Aylin Caliskan, Tadayoshi Kohno, Chirag Shah, Natasha Jaques·· 4 小时前AI 评分45
Persona Policies(PPol):为 LLM 智能体评估生成真实用户画像
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
AI 导读
研究团队提出 Persona Policies(PPol),一种即插即用的控制层,通过进化编码智能体自动发现用户画像生成程序,让 LLM 用户模拟器产生更真实的行为差异,同时保留原任务目标。
正文
Abstract:Large Language Model (LLM) agents are increasingly deployed in settings where they interact with diverse users, including those who are unclear, impatient, or reluctant to share information. However, collecting real interaction data at scale remains expensive. The field has turned to LLM-based \emph{user simulators} as stand-ins, but these simulators inherit the behavior of their underlying models: cooperative and homogeneous. As a result, agents that appear strong in simulation often fail in real human interactions. To narrow this gap, we introduce Persona Policies (PPol), a plug-and-play control layer that induces realistic behavioral variation in user simulators while preserving original task goals. Rather than hand-crafting personas, we employ an evolutionary coding agent to discover persona generation programs optimized for human-likeness and behavioral coverage over real user conversations. The evolved program generates diverse, human-like personas for any task in the domain. Across 4 benchmarks--including $\tau^2$-bench Retail and Airline, ColBench, and WildChat--evolved PPol yield 28-72% absolute gains in fitness score over the baseline simulator. In blinded evaluations, annotators judged PPol users as 'human' 80.4% of the time, nearly 2x more than the baseline simulators. Training agents with PPol also improves real-world performance: our user study with live human-agent interactions showed that fine-tuning with our method boosted task success by +23% over default baselines. PPol thus offers a novel approach to strengthen simulator-based evaluation and training without changing underlying tasks.
| Comments: | Preprint under review |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2605.12894 [cs.AI] |
| (or arXiv:2605.12894v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2605.12894 arXiv-issued DOI via DataCite |
Submission history
From: Harshita Chopra [view email]
[v1]
Wed, 13 May 2026 02:16:51 UTC (2,309 KB)
[v2]
Wed, 7 Oct 2026 05:02:45 UTC (5,048 KB)
来源:arXiv:cs.CL · arxiv.org