跳到正文
arXiv:cs.CL· Zhi Rao, Yucheng Zhou, Qianran Sun, Yiqing Huang, Longcan Yuan, Jiayi Hou, Chengwen Yao, Lin Cheng, Donghui Sun, Xiaoxin Chen, Jun Wan·· 3 小时前

SignRAG:统一检索增强的无 Gloss 手语翻译框架

SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation

AI 导读

研究者提出 SignRAG 统一框架,结合分层预训练、目标域检索增强与检索感知强化微调,将无 Gloss 手语翻译适配到 decoder-only LLM。该方法在多个 SLT 基准上取得新 SOTA,并成为首个在 CSL-Daily 全部指标上超越 Gloss 监督方法的无 Gloss 方案。代码与不同规模模型已在 GitHub 开源。

正文

Authors:Zhi Rao, Yucheng Zhou, Qianran Sun, Yiqing Huang, Longcan Yuan, Jiayi Hou, Chengwen Yao, Lin Cheng, Donghui Sun, Xiaoxin Chen, Jun Wan

View PDF HTML (experimental)

Abstract:Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{this https URL}{GitHub}, together with models of different sizes to support future academic research.
Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as: arXiv:2610.11371 [cs.CL]
  (or arXiv:2610.11371v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11371

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Rao Zhi [view email]
[v1] Thu, 8 Oct 2026 07:03:24 UTC (498 KB)

来源:arXiv:cs.CL · arxiv.org