跳到正文
arXiv:cs.LG· Luka Ribar, Jeevan Bhoot, Douglas Orr·· 5 小时前AI 评分37

Llama-Mobile:面向 VLM 的高效 2.7-bit 量化框架

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

AI 导读

研究者提出 Llama-Mobile 框架,用模型自身生成训练数据、无需访问原始训练流程,并设计 2.7-bit 每参数格式以在 Arm CPU 上高效执行。该框架将 Llama 3.2 11B Vision Instruct 压缩至 3.7 GB(8-bit 激活),在标准视觉问答任务上保持较强性能。

正文

View PDF HTML (experimental)

Abstract:Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2608.21134 [cs.CV]
  (or arXiv:2608.21134v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2608.21134

arXiv-issued DOI via DataCite

Submission history

From: Douglas Orr [view email]
[v1] Fri, 21 Aug 2026 14:10:31 UTC (2,387 KB)
[v2] Fri, 2 Oct 2026 16:39:11 UTC (2,376 KB)

来源:arXiv:cs.LG · arxiv.org