MET:基於理論且具文化意識的多語言道德推理
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
July 13, 2026
作者: Ayoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu, Peter Railton, Lu Wang
cs.AI
摘要
語言模型日益被用於跨不同語言與文化脈絡的道德決策,然而現有研究在以下三方面忽略了多語言性:1)多語言評估基準依賴直接翻譯,未能適應文化特有項目;2)道德推理的推論階段方法依賴靜態且以英語為中心的腳本,且缺乏道德理論基礎;3)道德決策的訓練方法通常需要來自較強模型或人類標註者的昂貴監督。我們透過三項貢獻填補這些缺口。首先,我們提出MCLASH,一個多語言道德決策基準,用以捕捉跨語言的具體文化道德直覺與社會規範。其次,我們提出MET(基於理論推理的多語言倫理學),這是一個兩步驟提示方法,其基礎來自心理學與哲學領域專家策劃的理論根據:模型先選擇特定情境與文化的根據,再以使用者的母語對其進行推理。第三,我們提出MET-D(MET蒸餾),透過一個無需外部監督的自我蒸餾訓練階段來強化第二個步驟。MET-D在三個不同規模與家族(Qwen3-4B、Qwen3-8B、Gemma3-4B)的模型上,相對於基礎模型,在MCLASH上平均提升3.71個巨觀F1分數,在MMoralExceptQA上平均提升4.23分數,其中Qwen3-8B在馬來語上的MCLASH提升最高達12.94分。我們進一步揭示,MET-D平均提升母語推理62.13分數,且有益的根據在不同文化間系統性地有所差異。綜合而言,這些貢獻為文化對齊、理論基礎的多語言道德推理開闢了道路。
English
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.