跳到正文
arXiv:cs.AI· Jonathan Ivey, Aimee Liang, Arthur Y. S. Wang, Madeline Mandell, Ziang Xiao, Anjalie Field·· 4 小时前

InterviewPlayground:评估 AI 访谈员的模拟环境

InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers

AI 导读

研究者开发了 InterviewPlayground 模拟环境,用基于社会理论的模拟受访者评估 AI 访谈员,并生成 InterviewReportCard 打分报告。

正文

View PDF HTML (experimental)

Abstract:Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt to participant behaviors. To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory. Simulated studies in InterviewPlayground produce an InterviewReportCard, which assesses the performance of AI interviewers using a suite of validated measures. To test whether our simulation-based evaluations predict performance with human participants, we conduct 15 real qualitative studies with five AI interviewers, three interview topics, and 450 human participants and compare them to simulated studies in InterviewPlayground. We find that AI interviewer performance in InterviewPlayground predicts performance in human studies with an average Pearson correlation of 0.86 across 12 measures, and the simulated interactions from InterviewPlayground reproduce key findings from behavioral analysis of AI interviewers in the human studies. Together, these findings support the validity of InterviewPlayground in assessing AI interviewer performance and examining potential failure modes. Our work contributes a simulation environment for AI interviewers supported with empirical validation, and more broadly, a roadmap for future work to develop validated, simulation-based evaluations of conversational AI systems.
Comments: Preprint. 23 pages
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.12023 [cs.AI]
  (or arXiv:2610.12023v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12023

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jonathan Ivey [view email]
[v1] Thu, 8 Oct 2026 14:20:46 UTC (362 KB)

来源:arXiv:cs.AI · arxiv.org