跳到正文
arXiv:cs.LG· Nilushika Udayangania, Kishor Nandakishora, Marimuthu Palaniswami·· 4 小时前AI 评分25

MemKD:面向紧凑循环神经网络的内存保留知识蒸馏

Learning to Remember: Distilling Memory Retention for Compact Recurrent Neural Networks

AI 导读

研究人员提出知识蒸馏框架 MemKD,通过专门的损失函数捕捉教师与学生模型在时间序列子序列上的记忆保留差异,使学生模型有效模仿教师行为。实验显示 MemKD 显著优于现有 SOTA 知识蒸馏方法,并能在多种压缩级别下匹配教师模型性能,大幅减少参数量和内存占用且精度损失不显著,适用于可穿戴设备与边缘计算平台的实时时间序列分析。

正文

View PDF HTML (experimental)

Abstract:Deep learning models, particularly recurrent neural networks and their variants, such as long short-term memory, have significantly advanced time series analysis. These models capture complex, sequential patterns in time series, enabling real-time assessments. However, their high computational complexity and large model sizes pose challenges for deployment in resource-constrained environments, such as wearable devices and edge computing platforms. Knowledge Distillation (KD) offers a solution by transferring knowledge from a large, complex model (teacher) to a smaller, more efficient model (student), thereby retaining high performance while reducing computational demands. Current KD methods, originally designed for computer vision tasks, neglect the unique temporal dependencies and memory retention characteristics of time series models. To bridge this gap, we propose a novel KD framework termed Memory-Discrepancy Knowledge Distillation (MemKD). MemKD leverages a specialized loss function to capture memory retention discrepancies between the teacher and student models across subsequences within time series data, ensuring that the student model effectively mimics the teacher's behaviour. This approach facilitates the development of compact, high-performing recurrent neural networks suitable for real-time, time series analysis tasks. We provide additional experiments, in-depth theoretical analysis, and insights into the proposed framework across extended time series benchmarks. Our experiments demonstrate that MemKD significantly outperforms state-of-the-art KD methods. Additionally, we demonstrate that it can match the teacher model's performance across a wide range of compression levels, achieving notable reductions in parameter count and memory usage without a significant loss in accuracy.
Comments: Preprint
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.06942 [cs.LG]
  (or arXiv:2610.06942v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.06942

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nilushika Udayangani Hewa Dehigahawattage [view email]
[v1] Sat, 3 Oct 2026 09:19:59 UTC (1,266 KB)

来源:arXiv:cs.LG · arxiv.org