arXiv:cs.LG· Zijie Xu, Bingrui Guo, Yiding Sun, Yiting Dong, Zhile Yang, Zhaofei Yu·· 4 小时前AI 评分28
共模误差限制低时间步深度脉冲Q网络(CMC-DSQN)
Common-Mode Errors Limit Low-Timestep Deep Spiking Q-Networks
AI 导读
研究发现低时间步深度脉冲Q网络(DSQN)的性能下降主要源于各动作Q值共享的共模误差,该误差会通过自举目标损害时序差分学习。为此提出的CMC-DSQN用辅助ANN补偿SNN输出中的共模误差,推理时可完全移除辅助ANN以保证SNN能效。在Atari和MiniAtar上,T=2时性能超SOTA DSQN基线近20%,T=4时进一步超过ANN基线。
正文
Abstract:Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices. In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making. However, existing DSQNs often require multiple simulation timesteps for competitive performance, increasing computational and energy costs, whereas reducing the timesteps can cause substantial performance degradation. We investigate this degradation from the perspective of Q-value estimation errors. By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets. Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs. At inference, greedy action selection can be performed directly from the SNN outputs, allowing the auxiliary ANN to be completely removed and preserving the energy efficiency of SNNs. Extensive experiments on Atari and MiniAtar environments demonstrate substantial performance improvements under low-timestep settings. CMC-DSQN outperforms state-of-the-art DSQN baselines by nearly $20\%$ at $T=2$ and further surpasses the ANN baseline at $T=4$.
| Subjects: | Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07808 [cs.NE] |
| (or arXiv:2610.07808v1 [cs.NE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07808 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zijie Xu [view email]
[v1]
Tue, 6 Oct 2026 06:01:28 UTC (1,079 KB)
来源:arXiv:cs.LG · arxiv.org