ChatPaper.aiChatPaper

언어 체인 정렬: 교차 언어 랭킹 선호도 최적화

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

August 24, 2026
저자: Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim
cs.AI

초록

대규모 언어 모델의 정렬은 영어 중심의 고품질 선호 데이터에 크게 의존하며, 이로 인해 다른 언어에서는 최적에 못 미치는 성능이 나타나곤 한다. 본 논문에서는 영어의 강건한 선호 지식을 활용하여 목표 언어의 선호 정렬을 촉진하는 새로운 프레임워크인 교차언어 순위 선호 최적화(CRPO)를 제안한다. 우리는 목표 언어와 영어 사이의 병렬 선호 쌍 내에 계층적 구조를 설계하여 언어 내(intra-lingual) 및 언어 간(inter-lingual) 선호를 공동으로 최적화함으로써 언어 적응과 출력 품질을 향상시킨다. LambdaLoss 프레임워크를 기반으로 하는 CRPO는 이진 비교 기반 최적화를 넘어 여러 후보 응답에 걸친 상대적 순위 신호를 제공한다. 다양한 자원 규모의 다섯 언어에 걸친 실험에서 CRPO는 지시 수행 능력과 지식 활용 능력 모두에서 표준 접근법을 일관되게 능가한다. 특히 다양한 가중치 체계에서 관찰된 강건한 성능 향상은 다국어 설정에서 우리의 계층적 설계의 실증적 효과를 더욱 뒷받침한다. 또한, CRPO가 보상 마진과 바람직한 응답의 로그 확률을 모두 크게 개선하여 교차언어 정렬을 위한 보다 안정적인 선호 다양체에 기여한다는 점을 확인한다. 우리의 코드는 https://github.com/dltmddbs100/CRPO에서 확인할 수 있다.
English
The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-lingual Ranking Preference Optimization~(CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preferences, thereby enhancing language adaptation and output quality. Building on the LambdaLoss framework, CRPO goes beyond the binary comparison based optimization by providing a relative ranking signal across multiple candidate responses. Our experiments across five languages with varying resource scales demonstrate that CRPO consistently outperforms standard approaches in both instruction-following and knowledge utilization capability. Notably, the robust performance gains observed across various weighting schemes further validate the empirical effectiveness of our hierarchical design in a multilingual setup. Furthermore, our findings highlight that CRPO significantly improves both reward margins and the log-probability of desirable responses, contributing to a more stable preference manifold for cross-lingual alignment. Our code is available at https://github.com/dltmddbs100/CRPO.