跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Junhyuk Huh, Seoungbin Bae, Dabeen Lee·· 14 小时前AI 评分35

因果 logistic bandits 在反事实公平约束下的 Minimax 最优遗憾

Minimax Optimal Regret for Causal Logistic Bandits with Counterfactual Fairness

AI 导读

研究在反事实公平约束下的因果 logistic bandits,证明覆盖条件是必要的:无覆盖限制时,事实不可区分但最优公平动作不同的环境会导致 Ω(T) 期望联合损失。

正文

View PDF HTML (experimental)

Abstract:We study causal logistic bandits with counterfactual fairness constraints. The causal structure is given through known factual and counterfactual feature maps that share an unknown logistic reward parameter, but the learner observes only factual rewards. Consequently, the directions determining counterfactual feasibility need not be identifiable from the available feedback. The closest prior analyses either omit a coverage condition or impose a comparatively strong one, and do not establish matching lower bounds. We first show that some coverage condition is necessary: without a coverage-type restriction, factually indistinguishable environments with different optimal fair actions force $\Omega(T)$ expected joint loss. Under a weaker full-rank condition on the factual covariance pooled across actions, we identify a target-specific information scale $V_\star$ that measures the difficulty of estimating rewards and counterfactual effects from factual feedback. We construct worst-case families satisfying this condition on which every policy incurs expected joint loss $\Omega\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}\right)$. We also give an explore--then--exploit procedure tuned using $V_\star$ and an adaptive algorithm that does not require its value. Both algorithms achieve $\max\{R_T,V_T\}=\widetilde{O}\left(\left[V_\star\min\{\log K,d\}\right]^{1/3}T^{2/3}+\kappa d/\sigma_0^2\right)$, where $R_T$ is regret relative to the best fair action and $V_T$ denotes the cumulative stage-wise positive violations. Thus the upper and lower bounds match in their leading dependence on $T$, $V_\star$, and $\min\{\log K,d\}$, up to logarithmic factors.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01377 [cs.LG]
  (or arXiv:2610.01377v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01377

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Seoungbin Bae [view email]
[v1] Thu, 1 Oct 2026 09:44:44 UTC (1,365 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org