跳到正文
arXiv:cs.CL· Takanori Ashihara, Kohei Matsuura, Masato Mimura·· 10 小时前AI 评分29

HINTT 提交第二届 MLC-SLM 挑战赛:级联与统一方法在说话人分离与 ASR 上的对比

HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR

AI 导读

HINTT 团队提交第二届 MLC-SLM 挑战赛的系统,针对多语言说话人归因 ASR,最终选用级联方案,由微调后的 DiariZen 分离模型、微调后的 Qwen3-ASR 和基于 LLM 的生成式纠错组成。对比实验中,同等官方数据微调的 VibeVoice-ASR 统一模型显示级联系统在 Task 1 条件下更可靠。全部微调与模型选择仅用官方 MLC-SLM 数据,未使用外部数据或伪标签。

正文

View PDF HTML (experimental)

Abstract:This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
Cite as: arXiv:2610.08063 [eess.AS]
  (or arXiv:2610.08063v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2610.08063

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Takanori Ashihara [view email]
[v1] Tue, 6 Oct 2026 09:57:17 UTC (112 KB)

来源:arXiv:cs.CL · arxiv.org