arXiv:cs.AI· Xin Wang, Boyan Gao, Yibo Yang, David A. Clifton·· 4 小时前AI 评分37
Mental-R1:为心理健康评估对齐 LLM 推理
Mental-R1: Aligning LLM Reasoning for Mental Health Assessment
AI 导读
研究提出 Cognitive Relative Policy Optimization(CRPO)强化学习框架,通过分阶段熵正则化模拟人类从不确定到确定的认知转变,并基于认知评价理论形式化推理阶段。在 8 个心理健康数据集上,CRPO 的加权 F1-score 较最佳强化学习基线平均提升 10.4 个百分点。经 CRPO 训练的 Mental-R1 在推理密集型案例上优于现有大语言模型。
正文
Abstract:Mental health problems such as anxiety, depression, and suicide remain urgent global challenges, where timely and accurate assessment is critical for effective intervention. Recently, large language models have been explored for mental health assessment. However, existing general-purpose post-training methods do not align with the cognitive processes of human assessment, which may lead to unreliable reasoning outcomes. To bridge this gap, we propose Cognitive Relative Policy Optimization (CRPO), a reinforcement learning framework tailored for the mental health domain. CRPO extends group relative policy optimization by integrating stage-dependent uncertainty modeling into the policy optimization process. Specifically, we introduce a stage-wise entropy regularization mechanism that encourages broad exploration in early reasoning phases and progressively enforces confident decision-making in later stages, mimicking the human cognitive shift from uncertainty to certainty. In addition, inspired by cognitive appraisal theory, we formalize cognitive reasoning stages, thereby guiding theory-grounded interpretable inference. Experiments on 8 mental health datasets show that CRPO achieves an average improvement of 10.4 percentage points in weighted F1-score over the best reinforcement learning baseline. Furthermore, the CRPO-trained model Mental-R1 demonstrates clear advantages compared with existing large language models on reasoning-intensive cases, suggesting that CRPO enhances reasoning capabilities for mental health assessment.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.13176 [cs.AI] |
| (or arXiv:2606.13176v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.13176 arXiv-issued DOI via DataCite |
Submission history
From: Xin Wang [view email]
[v1]
Thu, 11 Jun 2026 10:44:50 UTC (411 KB)
[v2]
Thu, 1 Oct 2026 21:40:44 UTC (422 KB)
来源:arXiv:cs.AI · arxiv.org