「Cultivar: 汚染とローカライゼーション堅牢性を調査するための対照的かつロケール指向の翻訳ベンチマーク」
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
August 10, 2026
著者: Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
cs.AI
要旨
多言語翻訳ベンチマークは通常、英語で作成され他言語に翻訳され、評価単位として言語ペアを用いる。この設計は、時間の経過とともに汚染されやすく、ロケールや文化的な考慮を見落としがちである。そこで我々はソース対照評価を提唱し、それをFLORESのローカライズ版サブセットであるCultivarとして具体化する。これにより、ロケール固有の翻訳評価が可能になる。ローカライズされていない対応物と組み合わせると、性能の乖離によってデータ汚染とローカライズ堅牢性を調査できる。我々は32のオープンウェイトモデルを評価し、機械翻訳特化モデルは堅牢性が低く、一部のモデルはFLORESに過適合している可能性があり、またモデルは言語に関係なく、他のロケールのコンテンツよりも米国向けコンテンツをうまく翻訳する傾向があることを見いだした。
English
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.