跳到正文
arXiv:cs.LG· Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem Taghva (Department of Computer Science, University of Nevada, Las Vegas)·· 3 小时前AI 评分43

SyntaxBench:面向大语言模型字符级推理的统计诊断框架

SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

AI 导读

研究者推出 SyntaxBench,一个针对大语言模型字符级推理的诊断基准与统计评估框架,包含字符计数、字母包含、回文检测、编辑距离、最长字符串选择五项核心任务,以及更难的子串抽取压力测试 index_to_span。

正文

View PDF HTML (experimental)

Abstract:Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts.
We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%.
Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.
Comments: 32 pages, 17 figures. The first two authors contributed equally. The code will be released soon
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
ACM classes: I.2.7; I.2.6
Cite as: arXiv:2610.03329 [cs.CL]
  (or arXiv:2610.03329v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.03329

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mohsen Larni [view email]
[v1] Fri, 2 Oct 2026 14:02:56 UTC (428 KB)

来源:arXiv:cs.LG · arxiv.org