跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Luca Geminiani, Nadja Klein·· 15 小时前AI 评分33

IQS-BO:面向贝叶斯优化的上下文查询选择方法

IQS-BO: In-Context Query Selection for Bayesian Optimisation

AI 导读

IQS-BO 是一种基于 PFN 的贝叶斯优化方法,通过对合成先验进行监督学习来学习查询决策,单次前向传播即可预测每个候选点最大化目标函数的概率。该方法提出查询的成本仅为基于采集函数方法的一小部分,在合成与真实基准上匹配或超越高斯过程标准 BO 及现有上下文方法。作者还提出一种混合先验用于预训练 PFN,结合 GP 样本与具有扭曲输入、孤立窄最优或平台期的函数,可提升优化性能。

正文

View PDF HTML (experimental)

Abstract:Bayesian Optimisation (BO) is a powerful framework for the optimisation of expensive black-box functions, but typically requires refitting a surrogate and maximising an acquisition function at every evaluation step. In-context approaches based on Prior-data Fitted Networks (PFNs) amortise part of this cost by pre-training transformers on functions drawn from synthetic priors. PFNs4BO amortises the surrogate but still relies on a numerically maximised acquisition function, while FIBO performs BO fully in-context by sampling optimiser locations from a learned density, which fixes the decision rule and admits no surrogate. Learned acquisition functions score a finite candidate set with a trained network, but, lacking a label for the query, learn the score by reinforcement learning on previously solved tasks. We propose IQS-BO, a PFN that learns the query decision by supervised learning on synthetic priors. In a single forward pass, IQS-BO predicts the probability that each candidate maximises the objective over the set, and we show that the minimiser of its objective is the posterior probability of this event. The model can be pre-trained without a surrogate for fully in-context BO, or take the predictions of a fixed probabilistic surrogate as additional input, amortising only the decision step. Our method proposes queries at a fraction of the cost of acquisition-based methods, while either matching or outperforming standard BO with Gaussian processes (GPs) and available in-context methods on synthetic and real-world benchmarks. Finally, we propose a mixture prior for pre-training PFNs which combines samples from GPs with functions exhibiting warped inputs, isolated narrow optima, or plateaus that are poorly modeled by stationary kernels common in GP surrogates. We show that pre-training on this prior can lead to improved optimisation performance.
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.01269 [cs.LG]
  (or arXiv:2610.01269v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01269

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Luca Geminiani [view email]
[v1] Thu, 1 Oct 2026 08:09:08 UTC (675 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org