arXiv:cs.AI· Wuraola Oyewusi, Eliana Vasquez Osorio, Goran Nenadic, Gareth Price·· 6 小时前AI 评分31
OncoNoteBERT:面向真实世界门诊肿瘤笔记 NLP 的基础表征模型
OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes
AI 导读
研究团队基于英国290,026份门诊肿瘤笔记(21,564名肺癌与头颈癌患者)开发了肿瘤专用 BERT 编码器 OncoNoteBERT 与 OncoNote-RadBERT。
正文
Abstract:Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoders using a governed UK outpatient oncology corpus comprising 290,026 notes from 21,564 patients treated for lung and head-and-neck cancer. We compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, and OncoNoteBERT, trained from scratch with an oncology WordPiece tokenizer. Models were evaluated using masked language modelling loss and perplexity on the validation set, tokenizer fragmentation metrics, clinical term tokenisation, masked-token probes, and exploratory representation analysis. Both external encoders fit the oncology corpus poorly in zero-shot evaluation (perplexity 113.04 for RadBERT; 2035.03 for PathologyBERT), while continued pretraining produced the strongest fit (2.10 for OncoNote-RadBERT). OncoNoteBERT achieved perplexity 2.83 but produced the most efficient tokenisation, with lower subword fertility and shorter normalised sequence length. It also returned a clinically acceptable prediction for 12 of 13 masked-token probes, compared with 7 of 13 for OncoNote-RadBERT. This divergence between corpus-level fit and masked-token performance was partly attributable to tokenizer fragmentation rather than learned semantics alone. Both locally developed models represented the institutional placeholder as a single learnable token. These findings show that continued adaptation and bespoke tokenisation provide complementary benefits, and that representation-layer design matters before adjacent-domain encoders are applied to oncology NLP.
| Comments: | Accepted at the AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models Workshop at NeurIPS 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Quantitative Methods (q-bio.QM) |
| Cite as: | arXiv:2610.03829 [cs.CL] |
| (or arXiv:2610.03829v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03829 arXiv-issued DOI via DataCite |
Submission history
From: Wuraola Fisayo Oyewusi [view email]
[v1]
Fri, 2 Oct 2026 11:38:07 UTC (3,819 KB)
来源:arXiv:cs.AI · arxiv.org