跳到正文
arXiv:cs.LG· Lee Seung-woo, Bowen Qi, Kim Min-jun, Jang Won-young·· 4 小时前AI 评分33

SkillFormer:面向音频语言模型的技能分解适配方法

SkillFormer: Skill-Decomposed Adaptation for Audio Language Models

AI 导读

SkillFormer 将音频理解分解为技能专属的低秩适配器,并通过可学习路由器在推理时按问题动态组合,让音高比较与曲风分类调用不同参数。它仅增加不到 4% 的基座模型参数,无需改动音频编码器或语言主干,在 MMSU、MMAU-Pro 和 MMAR 三个基准、三种架构不同的模型上平均准确率提升 2.5 至 4.1 分。

正文

View PDF HTML (experimental)

Abstract:Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose \textbf{SkillFormer}, which decomposes audio understanding into skill-specific low-rank adapters and composes them at inference time through a learned router. The router examines the question to decide which adapters to activate and how much weight each should carry, so that a pitch query engages different parameters than a genre classification query. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, preventing the gradient conflicts that arise in standard multi-task optimization. SkillFormer adds fewer than 4\% of the base model's parameters and requires no changes to the audio encoder or language backbone. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, it raises the average accuracy by 2.5 to 4.1 points, with balanced gains across perception, reasoning, and semantic subcategories.
Subjects: Sound (cs.SD); Machine Learning (cs.LG)
Cite as: arXiv:2610.07533 [cs.SD]
  (or arXiv:2610.07533v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.07533

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bowen Qi [view email]
[v1] Tue, 6 Oct 2026 00:00:21 UTC (24 KB)

来源:arXiv:cs.LG · arxiv.org