跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Bente Hinkenhuis, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag·· 14 小时前AI 评分27

多实例强化学习系统如何整合公平性与可解释性

Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System

AI 导读

一项研究提出结合 RL-MIL、对抗去偏与偏好条件超网络的多目标框架,用于学生风险预测,其中 MIL 将每名学生表示为弱标注交互的 bag,RL 智能体挑选信息量大的实例供下游分类。RL-MIL 基线分类表现强劲,但两种超网络变体均出现模式崩溃:改变偏好权重几乎无法沿公平性-性能前沿系统性移动。失败与目标主导、条件机制梯度传播弱及动态生成参数间的交互有关,表明仅靠偏好条件无法保证可控的多目标行为。

正文

View PDF HTML (experimental)

Abstract:Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of unfair predictions. This study investigates a multi-objective framework that combines reinforcement learning-based multiple instance learning (RL-MIL), adversarial debiasing, and preference-conditioned hypernetworks for student-at-risk prediction. MIL represents each student as a bag of weakly labeled interactions, while an RL agent selects informative instances for downstream classification. Two hypernetwork variants are evaluated to determine whether a user-defined preference scalar can continuously control the trade-off between predictive performance and Equalized Odds. The underlying RL-MIL baseline achieves strong classification performance, but both hypernetwork extensions exhibit mode collapse: changing the preference weight produces little systematic movement along the intended fairness-performance frontier. The failure is associated with objective dominance, weak gradient propagation through the conditioning mechanism, and interactions between dynamically generated parameters. The results show that fairness objectives can be incorporated into an interpretable RL-MIL pipeline, but preference conditioning alone does not guarantee controllable multi-objective behavior. Robust fair RL-MIL therefore requires explicit mechanisms for gradient balancing, objective separation, and stability analysis.
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2610.00035 [cs.LG]
  (or arXiv:2610.00035v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00035

arXiv-issued DOI via DataCite

Submission history

From: Seyed Sahand Mohammadi Ziabari [view email]
[v1] Wed, 2 Sep 2026 17:30:18 UTC (637 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org