跳到正文
arXiv:cs.CL· David Kletz, Sandra Mitrovi\'c, Ljiljana Dolami\'c, Fabio Rinaldi·· 6 小时前AI 评分35

语言模型是否具备文字系统感知能力?

Are Language Models Script-Aware?

AI 导读

一项被 AACL-IJCNLP 2026 接收的研究测试了 SLM 与 LLM 在多文字系统语言下的文字系统知识,即模型是否会随输入调整输出文字,以及能否按指令生成指定文字。测试模型在拉丁字母上的保真度均超过 98%,并能高频遵循文字指令;LLM 整体得分高于 SLM,在非标准文字组合上同样如此。

正文

View PDF HTML (experimental)

Abstract:Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.
Comments: Accepted to AACL-IJCNLP 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.08037 [cs.CL]
  (or arXiv:2610.08037v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08037

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: David Kletz [view email]
[v1] Tue, 6 Oct 2026 09:31:04 UTC (2,336 KB)

来源:arXiv:cs.CL · arxiv.org