arXiv:cs.LG· Geoffroy Morlat, Marceau Nahon, Augustin Chartouny, Raja Chatila, Ismael T. Freire, Mehdi Khamassi·· 6 小时前AI 评分35
COMETH:用概率聚类与大语言模型从人类数据中学习可解释的道德语境
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
AI 导读
研究者提出 COMETH 框架,将概率语境学习器与 LLM 语义抽象、人类道德评价结合,建模语境如何影响模糊行为的可接受度。该框架构建了覆盖 6 类核心行为、3 条道德规则(不杀、不欺、不违法)的 300 个场景数据集,并收集 N=101 参与者的 Blame/Neutral/Support 三元判断。
正文
Abstract:A key question in current AI alignment research is how to make AI algorithms learn moral values. Because human morality is highly context-dependent, actions are judged not only by their outcomes but by the context in which they occur. We present COMETH (Contextual Organization of Moral Evaluation from Textual Human inputs), a framework that integrates a probabilistic context learner with LLM-based semantic abstraction and human moral evaluations to model how context shapes the acceptability of ambiguous actions. We curate an empirically grounded dataset of 300 scenarios across six core actions relative to three moral rules (violating "Do not kill", "Do not deceive", and "Do not break the law") and collect ternary judgments (Blame/Neutral/Support) from N=101 participants. A preprocessing pipeline standardizes actions via an LLM filter and MiniLM embeddings with K-means, producing robust, reproducible core-action clusters. COMETH then learns action-specific moral contexts by clustering scenarios online from human judgment distributions using principled divergence criteria. To generalize and explain predictions, a Generalization module extracts concise, non-evaluative binary contextual features and learns feature weights in a transparent likelihood-based model. Empirically, COMETH roughly doubles alignment score with most human judgments relative to end-to-end LLM prompting (60% vs. 30% on average), while revealing which contextual features drive its predictions. The contributions are: (i) an empirically grounded moral-context dataset, (ii) a reproducible pipeline combining human judgments with model-based context learning and LLM semantics, and (iii) a more interpretable alternative than end-to-end LLMs for context-sensitive moral prediction and explanation.
| Comments: | 8 pages, 5 figures, 4 tables, +15 pages of Appendix |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2512.21439 [cs.CL] |
| (or arXiv:2512.21439v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2512.21439 arXiv-issued DOI via DataCite |
Submission history
From: Ismael Tito Freire González [view email]
[v1]
Wed, 24 Dec 2025 22:16:04 UTC (4,727 KB)
[v2]
Thu, 1 Oct 2026 23:02:56 UTC (4,736 KB)
来源:arXiv:cs.LG · arxiv.org