跳到正文
arXiv:cs.LG· Junhao Zhao, David Michael Simberg, Jacob Kang, Colin Connor Kurniawan, Nan Xu·· 4 小时前AI 评分41

Internal-DW:长时程自回归预测中远距离梯度未必可靠,需可靠性加权信用分配

Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting

AI 导读

针对长时程自回归预测中远距离梯度未必可靠的问题,研究者提出仅作用于反向传播的 Internal Dual-Wiener 路由(Internal-DW),在保留完整前向 rollout 与各步损失的同时对内部梯度路径做可靠性加权。

正文

View PDF HTML (experimental)

Abstract:In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate the update while amplifying predictable signal and unpredictable innovation together; a large distant gradient therefore need not carry reliable learning signal. Motivated by this, we introduce Internal Dual-Wiener routing (Internal-DW), a backward-only intervention that preserves the full forward rollout and all step losses while reliability-weighting internal gradient routes. At each residual block, we derive bounded Wiener gains for the identity and nonlinear routes that balance preserving predictable learning signal against suppressing unpredictable variation, and estimate them from route-level gradient statistics and an explicit noise model. In a controlled system with known gradient signal-to-noise ratio (SNR), we show that distant gradients can grow even as their SNR falls, and that Internal-DW reduces error in recovering predictable gradient signals and improves forecasting. On four history-dominated, weak-drive testbeds, Internal-DW reduces forecast error by 5.2%-13.8% relative to full BPTT, outperforms gradient clipping and Jacobian regularization on three testbeds, with similar performance on shear flow, and outperforms validation-selected truncated BPTT (TBPTT) on three. It also extends or preserves the fitted optimal training-horizon range across these four testbeds. Across benchmarks, the current Internal-DW estimator has a clear applicability boundary: its benefit diminishes or reverses when usable history is limited or when the selected sampler fails to represent dominant drive-dependent variation. The results show that retaining long-horizon supervision does not require trusting every backward contribution equally.
Comments: 38 pages, 9 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.12890 [cs.LG]
  (or arXiv:2609.12890v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.12890

arXiv-issued DOI via DataCite

Submission history

From: Junhao Zhao [view email]
[v1] Fri, 11 Sep 2026 14:18:59 UTC (678 KB)
[v2] Fri, 25 Sep 2026 06:46:40 UTC (807 KB)
[v3] Tue, 6 Oct 2026 21:35:19 UTC (805 KB)

来源:arXiv:cs.LG · arxiv.org