arXiv:cs.LG· Luka Ribar, Jeevan Bhoot, Douglas Orr·· 5 小时前AI 评分37
Llama-Mobile:面向 VLM 的高效 2.7-bit 量化框架
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
AI 导读
研究者提出 Llama-Mobile 框架,用模型自身生成训练数据、无需访问原始训练流程,并设计 2.7-bit 每参数格式以在 Arm CPU 上高效执行。该框架将 Llama 3.2 11B Vision Instruct 压缩至 3.7 GB(8-bit 激活),在标准视觉问答任务上保持较强性能。
正文
Abstract:Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2608.21134 [cs.CV] |
| (or arXiv:2608.21134v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.21134 arXiv-issued DOI via DataCite |
Submission history
From: Douglas Orr [view email]
[v1]
Fri, 21 Aug 2026 14:10:31 UTC (2,387 KB)
[v2]
Fri, 2 Oct 2026 16:39:11 UTC (2,376 KB)
来源:arXiv:cs.LG · arxiv.org