跳到正文
arXiv:cs.LG· Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna·· 4 小时前AI 评分32

AccentCL:抗类别不平衡与跨语料域偏移的英语口音分类增量学习框架

AccentCL: Robust Accent Classification with Incremental Expansion

AI 导读

AccentCL 是一个面向英语口音分类的类别增量学习框架,基于冻结的 Whisper-Large-v3 编码器提取多层表征,结合不平衡感知交叉熵损失与域均值对齐损失,并通过基于回放的持续学习扩展标签空间。

正文

View PDF HTML (experimental)

Abstract:Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
Comments: Published in Proceedings of IEEE Spoken Language Technology Workshop (SLT) 2026
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2610.07426 [cs.CL]
  (or arXiv:2610.07426v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07426

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mu-Ruei Tseng [view email]
[v1] Mon, 5 Oct 2026 21:34:23 UTC (660 KB)

来源:arXiv:cs.LG · arxiv.org