跳到正文
arXiv:cs.CL· Shuqing Shi, Ziyan Wang, Milind Tambe, Yali Du·· 5 小时前AI 评分36

HARP:通过显式人格推理实现异构偏好下的 LLM 编排

Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference

AI 导读

研究者提出 HARP(Heterogeneous-preference Agent oRchestration via Preference inference)框架,将 LLM 编排中对各智能体偏好的信念从提示词中移出,改为对每个智能体在有限候选偏好集上维护数值后验,并用 Bayes 规则闭式更新,语言模型只提供动作和逐候选似然。

正文

View PDF HTML (experimental)

Abstract:LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pursues but does not reveal. Inferring such hidden preferences from behavior has been a subject of long-standing research in game theory and multi-agent systems. The core challenge lies in maintaining a belief over every agent's preference and updating it from the agents' observed actions. Existing LLM orchestrators carry that belief as prompt text with no explicit update rule. This lets early errors persist and propagate rather than be corrected. We therefore propose \textbf{HARP} (Heterogeneous-preference Agent oRchestration via Preference inference), a novel framework that moves the belief out of the prompt. Specifically, HARP maintains one numeric posterior per agent over a finite set of candidate preferences and updates it in closed form by Bayes' rule. The language model supplies only actions and per-candidate likelihoods, so estimation is decoupled from its reasoning. We prove that HARP attains the same $\tilde O(\sqrt K)$ Bayesian regret as explicit joint inference when the factorization is exact. Furthermore, HARP\textsuperscript{+} augments planning with a bonus for actions that distinguish the candidates, so inference continues even when the optimal action is uninformative. Empirical results on three substrates, ranging from payoffs the preferences fully determine, through payoffs that depend on more than them, to scales where explicit joint inference is infeasible, demonstrate that HARP\textsuperscript{+} is the strongest non-oracle method across the class our theory identifies.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.07587 [cs.CL]
  (or arXiv:2610.07587v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07587

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shuqing Shi [view email]
[v1] Tue, 6 Oct 2026 01:24:56 UTC (1,047 KB)

来源:arXiv:cs.CL · arxiv.org