跳到正文
arXiv:cs.AI· Dylan Zongmin Liu·· 4 小时前AI 评分56

SovereignNegotiation-Bench:评估个人AI Agent在隐私、同意与机构压力下的代理谈判

SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure

AI 导读

论文提出 SovereignNegotiation-Bench,将代理法中的忠诚、服从授权、保密、坦率与勤勉五项义务转化为对对话日志的确定性检查,含 1,764 个配对场景(18 个消费与点对点领域、7 种对手策略)。

正文

View PDF HTML (experimental)

Abstract:Personal AI agents are beginning to negotiate for people, from refunds and bills to deposits and sales. A human agent in that position is judged by the duties owed to the principal, not by whether a deal was struck. We introduce SovereignNegotiation-Bench, a controlled benchmark that operationalizes five such duties from agency law--loyalty, obedience to actual authority, confidentiality, candor and diligence--as deterministic checks on episode logs; the first three enter a single headline metric. The benchmark contains 1,764 paired scenarios (252 situations in 18 consumer and peer-to-peer domains, each under 7 counterparty tactics). The counterparty's economics are a fixed function of the agent's structured actions and of the disclosures detected in its messages, so outcomes are comparable across agents and a disclosed limit has a measurable, causal price. A simulated principal grants or withholds consent and tightens its mandate midepisode. Rule-based agents show that the benchmark is solvable from the observable state (92% faithful success) and that a single disclosing sentence erases the entire negotiated surplus (0.70 to 0.00). Across 17 open-weight models, faithful success ranges from 6% to 75%; models disclose the principal's reservation value in 2-80% of episodes, agree or share a protected document without a required approval in 2-23%, and follow an instruction injected into the counterparty's message in 5-57% of injection episodes. Deal rate ranks models much like faithful success does, but it does not certify individual agreements: pooled over models, 48% of the agreements breach at least one duty (18-96% per model). Within families, faithful success tends to rise with size, but no size trend is significant, and on the model we test, neither prompting nor a code-level guard raises faithful success substantially. Code, scenarios and all episode logs will be released.
Subjects: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
Cite as: arXiv:2607.02814 [cs.MA]
  (or arXiv:2607.02814v2 [cs.MA] for this version)
  https://doi.org/10.48550/arXiv.2607.02814

arXiv-issued DOI via DataCite

Submission history

From: Zongmin Liu Dr. [view email]
[v1] Thu, 2 Jul 2026 23:03:15 UTC (52 KB)
[v2] Thu, 1 Oct 2026 18:37:42 UTC (91 KB)

来源:arXiv:cs.AI · arxiv.org