MET: 理論に基づいた文化を考慮した多言語道徳推論
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
July 13, 2026
著者: Ayoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu, Peter Railton, Lu Wang
cs.AI
要旨
言語モデルは、多様な言語・文化的文脈における道徳的意思決定にますます活用されているが、既存研究では以下の3点において多言語性が見落とされている:1)多言語評価ベンチマークが直訳に依存しており、文化固有の要素に適応できていない、2)道徳推論のための推論時手法が静的で英語中心の枠組みに頼り、道徳理論に基づいていない、3)道徳的意思決定のための訓練手法は、通常、より強力なモデルや人間によるアノテーションからの高コストな教師信号を必要とする。我々は3つの貢献を通じてこれらのギャップに対処する。第一に、言語横断的に文化に位置づけられた道徳的直感と社会的規範を捉える多言語道徳的意思決定ベンチマークMCLASHを導入する。第二に、心理学と哲学から得られた専門家厳選の理論に基づく根拠を活用した2段階プロンプト手法MET(理論に基づく多言語倫理)を提案する:モデルはまず状況・文化固有の根拠を選択し、次にユーザーの母語でそれらについて推論を行う。第三に、外部からの教師信号を一切必要としない自己蒸留訓練段階を通じて第二段階を強化するMET-D(MET蒸留)を導入する。MET-Dは、異なるサイズ・系統の3モデル(Qwen3-4B、Qwen3-8B、Gemma3-4B)すべてにおいて、ベースモデルと比較してマクロF1を平均でMCLASHでは3.71ポイント、MMoralExceptQAでは4.23ポイント向上させ、Qwen3-8Bのマレー語ではMCLASHで最大12.94ポイントの改善を達成した。さらに、MET-Dが母語による推論を平均62.13ポイント増加させること、有益な根拠が文化によって系統的に異なることを明らかにした。これらの貢献は、文化に適合し理論に基づいた多言語道徳推論への道を開くものである。
English
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.