arXiv:cs.CL· Marek \v{S}uppa, Ivan Vykopal, Andrej Ridzik, Kristi\'an Sopkovi\v{c}, Nat\'alia K\v{n}a\v{z}ekov\'a, Jaroslav Kop\v{c}an, Miroslav Bl\v{s}t\'ak, Vikt\'oria Ondrejov\'a, Daniel Hl\'adek, Michal Gregor, Martin Tamajka, Mari\'an \v{S}imko·· 3 小时前AI 评分38
sk-bench:面向斯洛伐克语大语言模型评测的原生优先基准
sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak
AI 导读
研究者发布 sk-bench,一个原生优先的斯洛伐克语基准,含 30 个数据集、33 个评分任务变体,覆盖十类技能,11 项资源为首次用于生成式 LLM 评测。
正文
Authors:Marek Šuppa, Ivan Vykopal, Andrej Ridzik, Kristián Sopkovič, Natália Kňažeková, Jaroslav Kopčan, Miroslav Blšták, Viktória Ondrejová, Daniel Hládek, Michal Gregor, Martin Tamajka, Marián Šimko
Abstract:Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($\rho\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($\rho=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at this https URL
| Comments: | Accepted to EMNLP 2026 Main |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2610.09152 [cs.CL] |
| (or arXiv:2610.09152v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09152 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Marek Šuppa [view email]
[v1]
Tue, 6 Oct 2026 21:51:22 UTC (236 KB)
来源:arXiv:cs.CL · arxiv.org