跳到正文
arXiv:cs.LG· Maria Zafar, Souhail Bakkali, Rejwanul Haque·· 4 小时前AI 评分29

生物医学领域神经机器翻译的模型压缩研究

Investigating Model Compression for Neural Machine Translation in the Biomedical Domain

AI 导读

研究将知识蒸馏与量化结合用于法语到英语的生物医学翻译,得到一个协同蒸馏并量化的学生模型,相比基线模型体积缩小 69%、推理速度提升 98.21%、CO2 排放减少 98.46%,且翻译质量未下降。该工作对比了多种微调策略以适配这一术语专业、平行语料稀缺的领域。

正文

View PDF HTML (experimental)

Abstract:Large-scale pretrained transformer models have achieved state-of-the-art performance across diverse machine translation tasks, including multilingual settings. Knowledge distillation has emerged as a sustainable approach for model compression, transferring knowledge from large teacher models to smaller, more efficient student models. Similarly, quantization, which reduces the numerical precision of model weights and activations (e.g., from 32-bit to 8-bit representations) is widely used to accelerate inference, enabling models to run several times faster during deployment. However, both techniques face limitations when applied to specialized domain data, particularly under low-resource conditions. In knowledge distillation, the effectiveness of transfer is often constrained by the scarcity of domain-specific parallel data, while quantization can lead to performance degradation as bit precision decreases. In this work, we investigate the combined application of knowledge distillation and quantization for French-to-English biomedical translation, a domain characterized by specialized terminology and limited parallel resources. We develop and compare multiple fine-tuning strategies to adapt compressed student models to this challenging setting. Our experiments demonstrate that a collaboratively distilled and quantized student model achieves a 69% reduction in size, a 98.21% increase in inference speed, and a 98.46% reduction in CO2 emissions compared to the original baseline all without sacrificing translation quality. These results indicate that jointly optimized compression techniques can yield efficient, high-performance models suitable for translation service providers operating under resource constraints.
Comments: Accepted at AICS 2025
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
MSC classes: 68T50
ACM classes: I.2.7; I.2.6
Cite as: arXiv:2610.07032 [cs.CL]
  (or arXiv:2610.07032v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07032

arXiv-issued DOI via DataCite (pending registration)

Related DOI: https://doi.org/10.5281/zenodo.19135901

DOI(s) linking to related resources

Submission history

From: Souhail Bakkali [view email]
[v1] Sun, 4 Oct 2026 19:01:38 UTC (45 KB)

来源:arXiv:cs.LG · arxiv.org