arXiv:cs.AI· Haoyue Liu, Zhichao Wang, Huanyu Yan, Xiaoying Tang·· 3 小时前
BudgetAPO:如何在紧预算下优化提示词?噪声自适应评估方案
How Should a Prompt Optimizer Spend a Tight Budget? BudgetAPO with Noise-Adaptive Evaluation
AI 导读
BudgetAPO 是一种面向紧预算场景的单阶段提示词优化器,通过噪声自适应评估切片、固定切片配对比较和反思算子,在七个基准和五个被测模型上全部排名第一,并通过 Holm 校正配对检验击败所有基线。
正文
Abstract:Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under tight budgets they fail in two ways: multi-stage pipelines can exhaust the budget and return the seed prompt unchanged, while single-stage methods compare candidates on fixed-size minibatches, regardless of each task's noise. As a remedy, we introduce BudgetAPO, a single-stage optimizer for the tight-budget regime. BudgetAPO incorporates (1) a noise-adaptive rule that sizes the evaluation slice to each task's noise, measured by a short probe; (2) a fixed slice that turns every accept/reject decision into a paired comparison; and (3) a reflective operator that rewrites reasoning strategy and output format jointly. Extensive results across seven benchmarks and five subject models demonstrate that BudgetAPO ranks first on every subject and beats every baseline under Holm-corrected paired tests, while returning the seed in 13% of runs at 250 calls against 86% for GEPA. On GPT-OSS-20B, GEPA needs 5 times as many calls to match \method's 100-call score.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.05671 [cs.AI] |
| (or arXiv:2610.05671v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.05671 arXiv-issued DOI via DataCite |
Submission history
From: Haoyue Liu [view email]
[v1]
Mon, 5 Oct 2026 01:34:14 UTC (6,677 KB)
[v2]
Thu, 8 Oct 2026 14:15:37 UTC (6,671 KB)
来源:arXiv:cs.AI · arxiv.org