Cultivar: 오염 및 현지화 견고성 연구를 위한 대조적·지역 지향적 번역 벤치마크
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
August 10, 2026
저자: Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
cs.AI
초록
다국어 번역 벤치마크는 일반적으로 영어로 작성된 원문을 다른 언어로 번역하여 언어 쌍을 평가 단위로 취급하는데, 이러한 설계는 시간이 지남에 따라 데이터 오염에 취약하며 로케일 및 문화적 고려 사항을 간과한다. 따라서 우리는 소스 대조 평가(source-contrastive evaluation)를 제안하며, 이를 FLORES의 현지화 하위 집합인 Cultivar로 구현하여 로케일별 번역 평가를 가능하게 한다. 현지화되지 않은 대응 세트와 짝지어 비교하면 성능 차이를 통해 데이터 오염과 현지화 견고성을 탐지할 수 있다. 우리는 32개의 오픈 가중치 모델을 평가한 결과, 기계번역(MT) 특화 모델이 덜 견고하며, 일부 모델은 FLORES에 과적합되었을 가능성이 있고, 모델들은 언어와 무관하게 미국 콘텐츠를 다른 로케일의 콘텐츠보다 더 잘 번역하는 경향이 있음을 발견했다.
English
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.