跳到正文
arXiv:cs.CL· Roshan Balaji, Pavan Kumar S, Vasudev Gupta, Sreejith N, Keerthana Sridhar, Nirav Bhatt·· 3 小时前

BioBigBird:面向生物医学文本长距离依赖处理的稀疏注意力模型

BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text

AI 导读

BioBigBird 是一个在生物医学文献与临床数据上预训练的双向语言模型,采用稀疏注意力机制处理最长 4096 tokens 的序列,并通过多阶段训练降低大规模预训练语料噪声。模型引入多任务学习框架,联合优化命名实体识别与关系抽取,在 BLURB 基准上取得与 SOTA 模型高度接近的结果。模型已公开可用。

正文

View PDF HTML (experimental)

Abstract:While domain-specific Large Language Models (LLMs) have encoded vast biomedical knowledge, their limited context windows often hinder a deep understanding of nuanced relationships within and across texts. To address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies. BioBigBird leverages a sparse attention mechanism to process sequences up to 4096 tokens, and its training incorporates a multi-stage process to mitigate noise from the large-scale pre-training corpus. We further enhance its performance by employing a multi-task learning (MTL) framework that jointly optimizes for Named Entity Recognition and Relation Extraction. Comprehensive evaluations on the BLURB benchmark reveal that our MTL-enhanced BioBigBird achieves highly competitive results against state-of-the-art models. Our work contributes an effective methodology for developing powerful, long-context language models for specialized domains, demonstrating the value of extended sequence processing for complex text analysis. Our models are publicly available at this https URL.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.11430 [cs.CL]
  (or arXiv:2610.11430v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11430

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Roshan Balaji [view email]
[v1] Thu, 8 Oct 2026 07:54:12 UTC (2,787 KB)

来源:arXiv:cs.CL · arxiv.org