arXiv:cs.LG(机器学习,全量分类)· Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, Serena Singh, Sepehr Bayat, Keith Pardee·· 14 小时前AI 评分35
蛋白质语言模型的量化与高效适配分析:QLoRA 与 4-bit 量化评估
Analysis of Quantized and Efficiently Adapted Protein Language Models
AI 导读
研究评估了 4-bit 量化与 QLoRA 在 ESM-2、ESMC、ProtBERT、ProtT5、Ankh、Ankh3 和 Profluent-E1 上的表现,多数模型-任务组合保留了全量微调 90% 以上的性能,最大模型的峰值 GPU 显存节省接近 90%。
正文
Authors:Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, Serena Singh, Sepehr Bayat, Keith Pardee
Abstract:Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized. Results: We evaluated 4-bit quantization and low-rank adapter fine-tuning (QLoRA) across ESM-2, ESMC, ProtBERT, ProtT5, Ankh, Ankh3 and Profluent-E1. Across protein prediction tasks, many model-task pairs retained more than 90% of full fine-tuning performance. Peak GPU memory savings approached 90% for the largest models, although performance and efficiency varied by model, dataset and training configuration. QLoRA often preserved early-layer representations while inducing task-specific adaptations in middle and late layers, resembling full fine-tuning with smaller representational changes. Training speed and power effects were more varied. For unconditional generation with ProLLaMA, ProtGPT2, ProGen2, ProteinGLM and ESM3, 4-bit quantization largely preserved predicted structural and sequence-level properties, but token-level analysis revealed model-dependent shifts in autoregressive output distributions. Conclusion: QLoRA and 4-bit quantization reduce PLM computational requirements, particularly GPU memory usage. Our results support QLoRA as a first-pass strategy for memory limited adaptation, reserving full fine-tuning for challenging tasks, unstable architectures or low validation recovery. For generative PLMs, sequence-level and structural metrics should be complemented with distributional analysis, since downstream predictions alone may miss quantization-induced shifts. These approaches can broaden access to large-scale protein modelling while requiring model- and task-specific validation.
| Subjects: | Machine Learning (cs.LG); Quantitative Methods (q-bio.QM) |
| Cite as: | arXiv:2610.00665 [cs.LG] |
| (or arXiv:2610.00665v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00665 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ilan Zeisler [view email]
[v1]
Wed, 30 Sep 2026 20:02:41 UTC (6,873 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org