跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran·· 15 小时前AI 评分32

重新思考信息瓶颈:标签诱导划分下的结构化分解

Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions

AI 导读

研究提出一种基于标签诱导划分的双瓶颈信息瓶颈(IB)公式,用标准 KL 项控制全局信息容量,用条件 KL 项约束条件内信息,并将条件 KL 精确分解为条件内信息项与先验失配项。该方法采用单纯形结构条件先验,可提供可控的潜在几何并无缝集成到现有流程中。分类与分割实验显示,低数据分类任务增益最明显,密集预测基准上取得一致提升。

正文

View PDF HTML (experimental)

Abstract:Standard information bottleneck (IB) regularization constrains representations via a single scalar I(Z;X), implicitlytreating all information as homogeneous. However, a single global compression control couples label-relevant structurewith residual within-condition variation, rather than regulating their allocation independently, allowing nuisanceinformation to persist in learned representations. For example, in medical imaging applications, residual variation oftenstems from acquisition conditions, background factors, or subject-specific appearance. This issue becomes particularlypronounced in data-limited settings, where models tend to overfit such variation, hindering generalization. While existingregularization methods can stabilize training, control capacity, or shape representation geometry, they do not explicitlyseparate nuisance-like variation from task-supporting structure. To address this limitation, we revisit IB from a structuredperspective based on a label-induced partition, where condition-level structure and within-condition information playdistinct roles. This leads to a dual-bottleneck formulation: a standard KL term controls global information capacity, while aconditional KL term targets within-condition information. We show that the conditional KL admits an exact decompositioninto a within-condition information term and a prior-mismatch term, explaining its alignment with the design this http URL a simplex-structured conditional prior, the method provides controllable latent geometry and integrates seamlesslyinto existing pipelines. Experiments on classification and segmentation show the clearest gains in low-data classificationand consistent improvements across dense prediction benchmarks.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01175 [cs.LG]
  (or arXiv:2610.01175v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01175

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lu Han [view email]
[v1] Thu, 1 Oct 2026 06:50:36 UTC (15,957 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org