跳到正文
arXiv:cs.CL· Yiping Bai·· 3 小时前AI 评分31

小规模中文 BERT 预训练对比:MLM、WWM 与 MacBERT 策略

Tiny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies

AI 导读

一项受控实验在 870 万参数、4 层 256 维的小规模中文 BERT 上从头训练并比较 MLM、WWM、MacBERT 三种预训练策略,语料为中文维基百科 129 万句。

正文

View PDF HTML (experimental)

Abstract:Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Masking (WWM), and MacBERT-style replacement have focused primarily on base-scale models (>=110M parameters). This paper presents a controlled comparison of these three strategies on a tiny-scale Chinese BERT model (4 layers, 256 hidden dimensions, 8.7M parameters). Under identical architecture, corpus (1.29M sentences from Chinese Wikipedia), and hyperparameters, we train three models from scratch and evaluate them across five intrinsic dimensions: perplexity, MLM hit rate, semantic discrimination, grammatical judgment, and contextual sensitivity. At tiny scale, MLM achieves the best overall intrinsic performance (winning 3 of 5 dimensions), while WWM excels in both perplexity (1.27 vs. 2.10, a 39.5% improvement) and MLM hit rate (22% vs. 16%). Notably, MacBERT under a severely limited synonym dictionary (222 entries, 3.3% coverage) exhibits severe perplexity degradation (47.23, 22x higher than MLM), yielding a ranking (MLM > WWM >> MacBERT) that differs markedly from the established base-scale conclusion (MacBERT > WWM > MLM). We further identify a critical evaluation pitfall: MacBERT achieves the lowest training loss (2.17) yet the highest perplexity (47.23), revealing that training loss alone is unreliable under mixed replacement strategies. All models and corpus are publicly available at this https URL.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08879 [cs.CL]
  (or arXiv:2610.08879v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08879

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yiping Bai [view email]
[v1] Tue, 6 Oct 2026 07:30:05 UTC (10 KB)

来源:arXiv:cs.CL · arxiv.org