arXiv:cs.LG· Nur A Zarin Nishat, Jens Lehmann, Andrei Aioanei, Sahar Vahdati·· 4 小时前AI 评分43
小语言模型在抽象推理任务上的系统性研究
A Systematic Study of Small Language Models on Abstract Reasoning Tasks
AI 导读
研究在 ARC-TGI 基准上系统测试了小语言模型的抽象推理能力,覆盖 decoder-only、encoder-decoder 和 MoE 三类模型家族、超 1000 次监督微调运行。结果显示分布内准确率可以很高,但技能获取对优化敏感且在不同任务族间分布不均,分布外性能急剧下降,即使规则保留而网格尺度变化也会失效。注意力诊断在部分任务上呈现不同集中与上下文依赖特征,但未能确立通用因果机制。
正文
Abstract:Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08680 [cs.LG] |
| (or arXiv:2610.08680v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08680 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nur A Zarin Nishat [view email]
[v1]
Tue, 6 Oct 2026 17:01:19 UTC (1,423 KB)
来源:arXiv:cs.LG · arxiv.org