跳到正文
arXiv:cs.CL· Melik\c{s}ah T\"urker, A. Ebrar K{\i}z{\i}lo\u{g}lu, Onur G\"ung\"or, Susan \"Usk\"udarl{\i}·· 6 小时前AI 评分35

TabiBERT:面向土耳其语的大规模 ModernBERT 基础模型与统一基准

TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish

AI 导读

研究人员推出基于 ModernBERT 架构的土耳其语单语编码器 TabiBERT,从零预训练 1 万亿 token,语料来自 865.8 亿 token 的多领域数据集(网页文本 72%、科学出版物 19%、源代码 6%、数学内容 0.3%),支持 8192 token 上下文,为现有土耳其语 BERT 模型的 16 倍。

正文

View PDF HTML (experimental)

Abstract:The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models remain the standard tool for classification, tagging and retrieval, where contextual representations and low inference cost matter more than text generation, yet Turkish lacks a monolingual encoder trained from scratch with the advances consolidated in ModernBERT (rotary positional embeddings, FlashAttention, refined normalization). We introduce TabiBERT, a monolingual Turkish encoder based on the ModernBERT architecture, pretrained from scratch for one trillion tokens sampled from an 86.58B-token multi-domain corpus of web text (72%), scientific publications (19%), source code (6%) and mathematical content (0.3%). The model supports a context length of 8,192 tokens, sixteen times that of existing Turkish BERT models, and inherits the ModernBERT architecture's efficiency at long context. For rigorous and reproducible evaluation we introduce TabiBench, a benchmark of 27 datasets across eight task categories with standardized splits and evaluation protocols, summarized as a GLUE-style macro-average on a 0-100 scale. TabiBERT leads the Turkish models in five of eight categories and BERTurk, the previous best, in six of eight; the gains concentrate on question answering (+9.55 F1) and code retrieval (+2.41 NDCG@10), while the four short-text categories are near saturation. Its average of 77.28 exceeds BERTurk's 75.66; the multilingual mmBERT reaches 78.98 with twice the parameters and three times the training tokens, at 41% more tokens per Turkish input. We release model weights, training configurations and evaluation code as a transparent and reproducible foundation for future Turkish encoder research.
Comments: 40 pages, 2 figures, 16 tables
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2512.23065 [cs.CL]
  (or arXiv:2512.23065v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2512.23065

arXiv-issued DOI via DataCite

Submission history

From: Meliksah Turker [view email]
[v1] Sun, 28 Dec 2025 20:18:22 UTC (169 KB)
[v2] Thu, 1 Jan 2026 13:14:04 UTC (412 KB)
[v3] Mon, 5 Jan 2026 10:15:27 UTC (412 KB)
[v4] Tue, 6 Oct 2026 15:23:27 UTC (174 KB)

来源:arXiv:cs.CL · arxiv.org