arXiv:cs.CL· David Kletz, Sandra Mitrovi\'c, Ljiljana Dolami\'c, Fabio Rinaldi·· 6 小时前AI 评分35
语言模型是否具备文字系统感知能力?
Are Language Models Script-Aware?
AI 导读
一项被 AACL-IJCNLP 2026 接收的研究测试了 SLM 与 LLM 在多文字系统语言下的文字系统知识,即模型是否会随输入调整输出文字,以及能否按指令生成指定文字。测试模型在拉丁字母上的保真度均超过 98%,并能高频遵循文字指令;LLM 整体得分高于 SLM,在非标准文字组合上同样如此。
正文
Abstract:Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.
| Comments: | Accepted to AACL-IJCNLP 2026 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08037 [cs.CL] |
| (or arXiv:2610.08037v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08037 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: David Kletz [view email]
[v1]
Tue, 6 Oct 2026 09:31:04 UTC (2,336 KB)
来源:arXiv:cs.CL · arxiv.org