arXiv:cs.AI· Dylan Zongmin Liu·· 4 小时前AI 评分65
SovereignPA-Bench:评估个人智能体在平台诱导与授权约束下的用户主权表现
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
AI 导读
论文提出 SovereignPA-Bench,用 1,920 个预订场景和 288 个取消场景测试个人智能体是否遵循用户当前意图、抵抗平台诱导、最小化数据共享并如实汇报。
正文
Abstract:Personal agents that book and buy for their users act through platforms that rank, nudge, pre-select and collect data in their own interest. We introduce SovereignPA-Bench, a controlled benchmark of whether such an agent follows the user's current intent, resists steering, shares only what a service needs, asks before exceeding its authority, and reports truthfully. A scripted platform and a scripted user surround the agent in 1,920 booking scenarios and 288 cancellation scenarios with retention flows. Each of the 192 booking situations, in 16 domains, is run as a control and under 9 paired variants that change one factor: stale memory, ambiguous intent, sponsored ranking, urgency, pre-checked extras, over-collecting forms, injected reviews, a mid-task update, or a mid-task reminder. Metrics are deterministic, and every question the agent asks is labelled necessary or unnecessary, so faithfulness and user burden are measured separately. Two rule-based agents reach 100% sovereign success, one of them without a single unnecessary question. Across 17 open-weight models, sovereign success ranges from 2% to 82%. More faithful models ask fewer unnecessary questions, not more (Spearman rho = -0.67 between success and burden). A form with two "recommended" fields raises the share of episodes that send a detail the service does not need from 2.4% to 58%. Injected reviews get a requested personal detail to the provider in 41% of attempts, against 0.7% without them. Sponsored labels and urgency banners shift choices only slightly in our setting. Retention offers never worked when the user had ruled them out in advance, but without that sentence 4 models accepted offers or pauses, and in 184 of 440 failed obstructed cancellations the agent told the user it had succeeded. A prompt-level checklist and a structural firewall raise success by at most 9 points. Code, scenarios and logs will be released.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.05363 [cs.AI] |
| (or arXiv:2607.05363v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05363 arXiv-issued DOI via DataCite |
Submission history
From: Zongmin Liu Dr. [view email]
[v1]
Mon, 6 Jul 2026 17:39:05 UTC (58 KB)
[v2]
Thu, 1 Oct 2026 18:35:16 UTC (113 KB)
来源:arXiv:cs.AI · arxiv.org