arXiv:cs.CL· Mengzhe Geng, Jinxi Ji, Junhao Xu·· 6 小时前AI 评分35
Qwen2-Audio 量化案例研究:文本分数无法确立词汇非诊断性语音任务的性能
Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study
AI 导读
一项针对 Qwen2-Audio-7B-Instruct 的量化案例研究显示,文本输出分数无法反映语音任务中量化是否保持性能。研究以平均 6 bit 和 7 bit 的固定混合 4/8-bit 配置,在 508 条 FLEURS 英译德语音和 512 条 RAVDESS 情感片段上评估,两者相对 FP16 的 BLEU 和 chrF 差异区间均包含零。
正文
Abstract:Text-output scores alone do not show whether quantization preserves performance on speech tasks whose target labels cannot be recovered from the transcript. We evaluate fixed mixed 4/8-bit Qwen2-Audio-7B-Instruct allocations averaging 6 and 7 bits per parameter on 508 English-to-German FLEURS utterances and on 512 RAVDESS emotion clips from 16 speakers. The BLEU and chrF differences from half precision (FP16) have intervals that include zero for both allocations. On RAVDESS, the same two sentences occur equally often with every emotion label. The absolute accuracy differences from FP16 are -3.71% for 6 bit and -1.17% for 7 bit. The 6-bit speaker interval excludes zero and an exact two-sided sign-flip test gives p=0.0148; the 7-bit interval includes zero. Same-budget controls do not identify either selected allocation as best. This case study shows why translation scores and performance on tasks beyond the transcript need separate evaluation.
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2609.26823 [cs.SD] |
| (or arXiv:2609.26823v2 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2609.26823 arXiv-issued DOI via DataCite |
Submission history
From: Mengzhe Geng [view email]
[v1]
Sun, 20 Sep 2026 14:57:27 UTC (62 KB)
[v2]
Tue, 6 Oct 2026 04:45:34 UTC (69 KB)
来源:arXiv:cs.CL · arxiv.org