arXiv:cs.LG· Yuhe Hu·· 4 小时前AI 评分37
量化对工具故障恢复的影响随提示词与评估设计而变:Llama-3.1-8B-Instruct 与 Qwen2.5-7B-Instruct 的 8-bit/4-bit 对比研究
Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs
AI 导读
Llama-3.1-8B-Instruct 与 Qwen2.5-7B-Instruct 的 8-bit 和 4-bit 量化版本,在 20 个确定性工具使用任务与 5 组提示词上的工具故障恢复表现并不稳定。
正文
Abstract:Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated. We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts. The 8-bit-4-bit recovery comparison changes direction across prompts and evaluation targets. On tasks that both variants complete without faults under the same prompt, the difference ranges from 0 to +20.2 percentage points for Llama and from -50.0 to +35.0 points for Qwen. Full-pipeline point estimates favor 8-bit Llama under all five prompts, whereas the Qwen comparison changes direction across prompts. The evaluation target can also reverse the result. For Llama under one prompt, scoring each variant only on its own clean-passing tasks favors 4-bit by 17.5 points; scoring the same tasks for both variants gives no difference, while scoring the full pipeline favors 8-bit by 28.3 points. Executor leniency is a third such choice. Rescoring the same logs with strict output parsing, which 8-bit Llama violates far more often than 4-bit Llama under that prompt, turns that +28.3 into -15.0 while leaving Qwen essentially unchanged. These findings show that one prompt, one screened task set, and one scoring policy do not establish a stable conclusion about quantized-agent robustness. Evaluations should compare variants on matched tasks, report full-pipeline success for deployment decisions, state the scoring policy, and quantify uncertainty across tasks rather than injected fault sites.
| Comments: | Accepted at the NeurIPS 2026 Workshop on Small Language Models for Agentic Systems (SLM-Agents). 7 pages, 2 figures, 2 tables, plus appendix |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2610.07781 [cs.AI] |
| (or arXiv:2610.07781v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07781 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuhe Hu [view email]
[v1]
Tue, 6 Oct 2026 05:17:10 UTC (97 KB)
来源:arXiv:cs.LG · arxiv.org