跳到正文
arXiv:cs.LG· Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen, Yixin Liu, Zhen Peng, Bin Shi, Bo Dong, Chao Shen, Irwin King, Qinghua Zheng·· 4 小时前AI 评分42

DecepEval:评估 LLM 智能体欺骗行为的基准

DecepEval: A Benchmark for Evaluating Deception in LLM Agents

AI 导读

研究者推出 DecepEval 基准,含 3 类任务族、28 个专业场景共 1,532 个实例,用于评估 LLM 智能体的欺骗行为。该基准基于经典欺诈理论提出 LLM Deception Diamond 框架,刻画压力、激励、机会与冲突四类诱导欺骗的外部条件,并通过中性版与诱导版实例配对测量欺骗率变化。

正文

Authors:Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen, Yixin Liu, Zhen Peng, Bin Shi, Bo Dong, Chao Shen, Irwin King, Qinghua Zheng

View PDF HTML (experimental)

Abstract:As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors. Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates. DecepEval makes these vulnerabilities measurable, providing a shared benchmark for progress toward trustworthy artificial intelligence.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07967 [cs.LG]
  (or arXiv:2610.07967v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07967

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yiming Xu [view email]
[v1] Tue, 6 Oct 2026 08:33:43 UTC (1,358 KB)

来源:arXiv:cs.LG · arxiv.org