arXiv:cs.LG· Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna·· 4 小时前AI 评分32
AccentCL:抗类别不平衡与跨语料域偏移的英语口音分类增量学习框架
AccentCL: Robust Accent Classification with Incremental Expansion
AI 导读
AccentCL 是一个面向英语口音分类的类别增量学习框架,基于冻结的 Whisper-Large-v3 编码器提取多层表征,结合不平衡感知交叉熵损失与域均值对齐损失,并通过基于回放的持续学习扩展标签空间。
正文
Abstract:Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
| Comments: | Published in Proceedings of IEEE Spoken Language Technology Workshop (SLT) 2026 |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2610.07426 [cs.CL] |
| (or arXiv:2610.07426v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07426 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mu-Ruei Tseng [view email]
[v1]
Mon, 5 Oct 2026 21:34:23 UTC (660 KB)
来源:arXiv:cs.LG · arxiv.org