MET: 이론에 근거하고 문화를 고려한 다국어 도덕 추론
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
July 13, 2026
저자: Ayoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu, Peter Railton, Lu Wang
cs.AI
초록
언어 모델은 다양한 언어 및 문화적 맥락에서 도덕적 의사 결정에 점점 더 많이 활용되고 있지만, 기존 연구는 다음 세 가지 측면에서 다국어성을 간과하고 있다: 1) 다국어 평가 벤치마크는 직접 번역에 의존하여 문화 특화 항목을 적응시키지 못한다; 2) 도덕적 추론을 위한 추론 시 방법은 정적인 영어 중심의 스캐폴드(틀)에 의존하며 도덕 이론에 기반하지 않는다; 3) 도덕적 의사 결정을 위한 훈련 방법은 일반적으로 더 강력한 모델이나 인간 주석자로부터의 값비싼 감독을 필요로 한다. 우리는 세 가지 기여를 통해 이러한 격차를 해결한다. 첫째, 언어 전반에 걸쳐 문화적으로 자리 잡은 도덕적 직관과 사회적 규범을 포착하기 위해 다국어 도덕적 의사 결정 벤치마크인 MCLASH를 도입한다. 둘째, 심리학과 철학에서 추출된 전문가 선별, 이론 기반 근거 위에 구축된 두 단계 프롬프팅 기법인 MET(다국어 이론 기반 윤리 추론)를 제안한다: 모델은 먼저 상황 및 문화 특화 근거를 선택한 후, 사용자의 모국어로 이에 대해 추론한다. 셋째, 외부 감독이 필요 없는 자기 증류(self-distillation) 훈련 단계를 통해 두 번째 단계를 강화하는 MET-D(MET-증류)를 도입한다. MET-D는 세 가지 서로 다른 크기와 계열의 모델(Qwen3-4B, Qwen3-8B, Gemma3-4B) 모두에서 기본 모델 대비 거시 F1(macro-F1)을 평균 3.71점(MCLASH) 및 4.23점(MMoralExceptQA) 향상시켰으며, Qwen3-8B에서 말레이어의 경우 MCLASH에서 최대 12.94점의 향상을 보였다. 나아가 MET-D가 모국어 추론을 평균 62.13점 증가시키며, 유익한 근거가 문화에 따라 체계적으로 다르다는 점을 밝힌다. 이러한 기여들은 함께 문화에 부합하고 이론에 기반한 다국어 도덕적 추론을 위한 길을 열어준다.
English
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.