arXiv:cs.CL· Songtao Li, Yijia Zhang, Jianyuan Yuan, Shidi Zhang, Fengyu Zhang, Hongfei Lin·· 4 小时前AI 评分34
MITE:用多编程语言指令微调与集成方法提升生物医学命名实体识别
Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method
AI 导读
研究者提出 MITE,将生物医学命名实体识别(BioNER)重构为结构到结构的生成任务,把指令与实体输出统一表示为 Python、C++、Java 等代码格式,无需外部生物医学知识或额外标注即可获得结构多样的监督信号。
正文
Abstract:Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annotations as flat textual outputs, providing limited structural constraints for typed entity extraction. Second, high-quality biomedical annotations are limited, and learning from a single serialized output form may restrict structural diversity and reduce model robustness. Although external biomedical knowledge can be introduced to alleviate data scarcity, it often requires costly resource construction. To address these challenges, we propose MITE, a Multiple Programming Languages Instruction Tuning and Ensemble method for BioNER. MITE reformulates BioNER as a structure-to-structure generation task by representing both instructions and entity outputs in code-formatted representations. Specifically, each training instance is transformed into multiple programming-language formats, including Python, C++, and Java, while preserving the same underlying entity semantics. These language-specific representations provide structurally diverse supervision without requiring external biomedical knowledge or additional annotations. During inference, MITE aggregates predictions from different code formats through an entity-level voting strategy, reducing language-specific prediction variance and improving robustness. Experiments on six widely used BioNER datasets demonstrate that MITE consistently outperforms representative BERT-based and LLM-based baselines and exhibits strong cross-dataset generalization. Ablation and parameter analyses further verify the effectiveness and robustness of the proposed components.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02949 [cs.CL] |
| (or arXiv:2610.02949v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02949 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | 2026 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2026 |
Submission history
From: Songtao Li [view email]
[v1]
Fri, 2 Oct 2026 07:42:44 UTC (458 KB)
来源:arXiv:cs.CL · arxiv.org