跳到正文
arXiv:cs.CL· Mehrdad Ghassabi, Hamidreza Baradaran Kashani, Pedram Rostami, Sadra Hakim, Zahra Kazemi, Amirhossein Poursina, Milad Tavakoli, Audrina Ebrahimi·· 4 小时前AI 评分33

Gaokerena:面向消费级硬件的波斯语医疗小模型家族

Gaokerena: A Small Persian Medical Language Model Family

AI 导读

论文提出 Gaokerena,一个可在消费级硬件部署的紧凑型波斯语医疗语言模型家族。Gaokerena-V 用 5700 万新 token 训练,将翻译版医疗 MMLU 从 46.64% 提升至 49.31%;Gaokerena-R 结合 Chain-of-Thought 与两套 RLAIF 框架,以更小数据集取得 52.98% 的更高分数。

正文

View PDF HTML (experimental)

Abstract:The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low-resource languages like Persian significantly underserved. To address this gap, this paper introduces Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer-grade hardware. As a foundational step toward localized digital healthcare, we first present Gaokerena-V, developed by training a baseline model on a strategically selected subset of a newly curated 90-million-token Persian medical corpus (approximately 54 million tokens) together with 20,000 expert-vetted physician Q&A pairs (approximately 3 million tokens), for a total of 57 million new tokens. This training improved performance on a translated medical MMLU benchmark from 46.64% to 49.31%. Second, recognizing the critical demands of clinical reasoning, we developed Gaokerena-R by integrating a Chain-of-Thought approach with two novel Reinforcement Learning with AI Feedback (RLAIF) frameworks to optimize preference-based reasoning. Despite utilizing the same baseline architecture and a smaller dataset than Gaokerena-V, Gaokerena-R achieved a superior benchmark score of 52.98%. Furthermore, both models are equipped with custom-developed uncertainty heads that predict the models confidence in its responses based solely on internal hidden states. While these results demonstrate significant progress in Persian medical language modeling and proactive safety estimation, current performance levels remain insufficient for direct clinical application, highlighting the necessity for further research into robust knowledge acquisition and rigorous safety verification prior to real-world deployment.
Comments: 37 pages, 9 figures
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2608.00932 [cs.CL]
  (or arXiv:2608.00932v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2608.00932

arXiv-issued DOI via DataCite

Submission history

From: Mehrdad Ghassabi [view email]
[v1] Sun, 2 Aug 2026 02:17:10 UTC (499 KB)
[v2] Thu, 3 Sep 2026 07:49:09 UTC (503 KB)
[v3] Thu, 24 Sep 2026 09:52:22 UTC (503 KB)
[v4] Fri, 2 Oct 2026 15:39:40 UTC (503 KB)

来源:arXiv:cs.CL · arxiv.org