跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Cagri Temel·· 14 小时前AI 评分37

四种分类器生长方式,以及其中一种为何无法学习

Four Ways to Grow a Classifier and Why One of Them Cannot Learn

AI 导读

论文在统一协议下比较树结构与构造式分类器的四种生长决策,并证明最自然的软决策树加深方式无法学习:把每个叶子变成门、两个子节点继承父节点类别分布后,所有新门的梯度恒为零,门取 1/2 时两子节点梯度相同。五折交叉验证三个随机种子下,该构造在 Iris、Wine、Digits 上分别比同深度从头训练低 19.6、19.1、55.6 个准确率点;修复方法是对子节点做小幅随机扰动,其幅度几乎无关紧要。

正文

View PDF HTML (experimental)

Abstract:Constructive classifiers add structure while they train: a level to a tree, a unit to a hidden layer, a split at a leaf. This paper asks what each of four such growth decisions actually buys, measured under one fixed protocol in tree-structured and constructive models, and gives an exact diagnosis and a fix for the one that buys nothing.
The diagnosis concerns the most natural way to deepen a soft decision tree: turn every leaf into a gate whose two children inherit the parent's class distribution, so that the function is unchanged. I prove that this leaves the gradient of every new gate identically zero and, with the gate at 1/2, gives the two children identical gradients, so the added level can never learn. Unlike the symmetry that Net2Net breaks with noise or the saddle point that splitting steepest descent escapes with second-order information, first-order information here is not weak but absent. Over three seeds of five-fold cross-validation the construction loses 19.6 accuracy points on Iris, 19.1 on Wine and 55.6 on Digits against the same depth trained from scratch. The fix is a small random perturbation of the children, whose size barely matters. The practical rule is one line in a test: after adding parameters, assert that their gradient is nonzero.
The other three decisions each buy one thing. Fitting a new hidden unit to the residual error before installing it buys a smaller network on every dataset, though not a more accurate one, and on Digits it costs accuracy significantly. Splitting the leaf with the largest expected error buys sparsity, reaching 0.885 with 3.7 splits where a complete depth-six tree uses 63, but loses 4.3 points on a harder problem. Requiring statistical significance before a node receives a more expressive split buys nothing: the tree gets larger and less accurate.
Every number in the paper is inserted from the measurement script.
Comments: 10 pages, 4 tables. Code and measurement scripts: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00180 [cs.LG]
  (or arXiv:2610.00180v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00180

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Cagri Temel [view email]
[v1] Wed, 16 Sep 2026 22:00:33 UTC (12 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org