arXiv:cs.LG· Junjie Xiao, Huiwen Jia·· 7 小时前AI 评分36
大规模初始化下逻辑梯度下降轨迹逼近的理想路径
Ideal Paths for Approximating Logistic Gradient Descent Trajectories at Large Initialization
AI 导读
研究针对严格线性可分数据、初始规模为 R 的大初始化场景,用最小范数投影规则构造出一条由有限线性段组成的唯一连续理想路径,并证明固定步长 GD 轨迹除以 R 后随 R→∞ 一致收敛到该路径。该路径分负间隔修正与最小间隔增长两阶段,修正阶段与增长阶段的累积损失分别按 R² 和 R 归一化后收敛到显式极限,峰值评估损失在两端点损失均趋于零时仍可随 R 线性增长。
正文
Abstract:Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected by prolonged training, but this direction alone does not provide information regarding the intermediate behavior. We address this question through a geometric approximation of full-batch logistic gradient descent (GD) trajectories on strictly linearly separable data, with large initialization of scale $R$ motivated by prior training. From any limiting normalized initial position, we use minimum-norm projection rules to construct a unique continuous ideal path consisting of finitely many linear segments. The path has two stages: negative-margin correction followed by minimum-margin growth. We prove that, after an explicit two-stage time reparameterization, the fixed-step GD trajectory divided by $R$ converges uniformly to this path on every fixed parameter interval as $R\to\infty$. Further, our quantitative error bounds account for initialization perturbations and the transition between stages. This approximation provides asymptotic formulas for peak evaluation loss and cumulative training loss. In particular, peak evaluation loss can grow linearly in $R$ even when both endpoint losses tend to zero. The cumulative losses in the correction and margin-growth stages, normalized by $R^2$ and $R$, respectively, converge to explicit limits. Experiments on controlled geometries and fixed image features complement our theoretical results.
| Comments: | 42 pages, 5 figures, 6 tables |
| Subjects: | Machine Learning (cs.LG); Optimization and Control (math.OC) |
| Cite as: | arXiv:2610.04142 [cs.LG] |
| (or arXiv:2610.04142v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.04142 arXiv-issued DOI via DataCite |
Submission history
From: Huiwen Jia [view email]
[v1]
Fri, 2 Oct 2026 23:29:20 UTC (160 KB)
[v2]
Tue, 6 Oct 2026 06:01:15 UTC (160 KB)
来源:arXiv:cs.LG · arxiv.org