arXiv:cs.AI· Hanfei Liu, Shuchang Feng, Yanxia Chen, Changzeng Fu, Shiqi Zhao·· 3 小时前
Spora 如何重新权衡脉冲语言模型的时间编码与非线性计算
Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models
AI 导读
Spora 通过联合设计脉冲编码与注意力算子,缓解脉冲语言模型在时间窗口语义表达与非线性注意力计算之间的权衡。其二进制时间权重让 T 个脉冲可表示至多 T 比特的组合值,而脉冲计数读出仅 O(log₂T) 比特;UBS 用阈值与脉冲触发残差衰减生成非负整数编码,BBS 分离符号与幅值并学习有符号激活的缩放。
正文
Abstract:Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations. We introduce Spora, which jointly designs spike encodings and attention operators. Binary temporal weights let $T$ spikes represent compositional values with up to $T$ bits of capacity, compared with $O(\log_2 T)$ bits for spike-count readout. Unipolar Binary Spiking (UBS) uses thresholds and spike-triggered residual decay to produce non-negative integer codes; Bipolar Binary Spiking (BBS) separates sign and magnitude and learns a scale for signed activations. These representations support accumulation-and-shift dot products and integer-exponent mappings in attention. With four time steps, Spora achieves 76.6 average GLUE score and 44.1 CoLA MCC, improving over SpikeLM by 1.2 and 6.2 points, respectively. Extending BBS to six steps raises these scores to 78.2 and 47.4. Conditional-decay analysis, matched-budget activation-quantization comparisons, event-workload statistics, and fixed-point evaluation further characterize the connection between encoding fidelity and computational cost.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10933 [cs.LG] |
| (or arXiv:2610.10933v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10933 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hanfei Liu [view email]
[v1]
Wed, 7 Oct 2026 21:41:46 UTC (4,883 KB)
来源:arXiv:cs.AI · arxiv.org