arXiv:cs.CL· Zhi Rao, Yucheng Zhou, Qianran Sun, Yiqing Huang, Longcan Yuan, Jiayi Hou, Chengwen Yao, Lin Cheng, Donghui Sun, Xiaoxin Chen, Jun Wan·· 3 小时前
SignRAG:统一检索增强的无 Gloss 手语翻译框架
SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation
AI 导读
研究者提出 SignRAG 统一框架,结合分层预训练、目标域检索增强与检索感知强化微调,将无 Gloss 手语翻译适配到 decoder-only LLM。该方法在多个 SLT 基准上取得新 SOTA,并成为首个在 CSL-Daily 全部指标上超越 Gloss 监督方法的无 Gloss 方案。代码与不同规模模型已在 GitHub 开源。
正文
Authors:Zhi Rao, Yucheng Zhou, Qianran Sun, Yiqing Huang, Longcan Yuan, Jiayi Hou, Chengwen Yao, Lin Cheng, Donghui Sun, Xiaoxin Chen, Jun Wan
Abstract:Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{this https URL}{GitHub}, together with models of different sizes to support future academic research.
| Subjects: | Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM) |
| Cite as: | arXiv:2610.11371 [cs.CL] |
| (or arXiv:2610.11371v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11371 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rao Zhi [view email]
[v1]
Thu, 8 Oct 2026 07:03:24 UTC (498 KB)
来源:arXiv:cs.CL · arxiv.org