ChatPaper.aiChatPaper

MET:基于理论且具备文化意识的多语言道德推理

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

July 13, 2026
作者: Ayoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu, Peter Railton, Lu Wang
cs.AI

摘要

语言模型正越来越多地被用于跨语言和文化背景的道德决策,但现有研究在三个方面忽视了多语言性:1)多语言评估基准依赖直接翻译,未能适应文化特定内容;2)道德推理的推理时方法依赖于静态、以英语为中心的框架,且缺乏道德理论基础;3)道德决策的训练方法通常需要来自更强模型或人工标注者的昂贵监督。我们通过三项贡献填补这些空白。首先,我们提出了MCLASH,一个多语言道德决策基准,旨在捕获跨语言的文化嵌入道德直觉和社会规范。其次,我们提出了MET(基于理论推理的多语言伦理方法),这是一种两步提示方法,建立在来自心理学和哲学领域的专家策划、基于理论的依据之上:模型首先选择情境和特定文化的依据,然后用用户的母语进行推理。第三,我们提出了MET-D(MET蒸馏),该方法通过一个无需外部监督的自蒸馏训练阶段增强第二步。MET-D在三个不同规模和家族的模型(Qwen3-4B、Qwen3-8B、Gemma3-4B)上相比基础模型平均提高了MCLASH上的宏观F1 3.71分,在MMoralExceptQA上提高4.23分,其中在Qwen3-8B上马来语的MCLASH增益最高达到12.94分。我们进一步揭示,MET-D平均增加了62.13分的母语推理,并且有益依据在不同文化中存在系统性差异。这些贡献共同为文化对齐、理论奠基的多语言道德推理开辟了道路。
English
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.