Cultivar:一個探討污染與在地化穩健性之對比式區域導向翻譯評測基準
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
August 10, 2026
作者: Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
cs.AI
摘要
多語言翻譯基準通常以英語為來源、翻譯成其他語言,並以語言對作為評估單位——這種設計隨著時間推移容易受到資料污染,且忽略了地域與文化差異。因此,我們提倡來源對比式評估,並以 Cultivar——FLORES 的在地化子集——來具體實踐此一方法,實現針對特定地域的翻譯評估。當與非在地化的對應版本搭配使用時,效能差異可協助探查資料污染與在地化穩健性。我們評測了 32 個開放權重模型,結果發現專精於機器翻譯的模型穩健性較低,少數模型可能過度擬合 FLORES,且無論語言為何,模型翻譯美國內容的表現普遍優於其他地域的內容。
English
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.