ChatPaper.aiChatPaper

NOLLI:英語と韓国語の性能ギャップを診断するための難易度較正済みパズルベンチマーク

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

August 5, 2026
著者: Dasol Choi, Joonyong Park, Daegon Yu, Soo Yong Kim, Youngsook Song, Seunghyeok Hong
cs.AI

要旨

本稿では、韓国語の性能格差がどこで生じるかを診断するために設計された、手続き的に生成される英語-韓国語パズルベンチマークであるNOLLIを紹介する。NOLLIは15種類のパズル(25タスク、7,500項目)で構成され、すべてのインスタンスはシードから再生成可能であり、一意の解を持つことが検証され、決定的に採点される。我々は、「難しい」ことを「大きい」ことと同一視するのではなく、難易度を行動的に較正し、固定された参照モデルが目標精度帯域に入るまで各生成器を調整する。その3レベル設計は、対応する直接翻訳、ハングル字母(音節以下の文字)にわたる文字体系適応、そして韓国の文化または正書法に基づく韓国語のみのタスクを組み合わせたものである。我々は、フロンティアモデル、オープンウェイトモデル、韓国開発モデルを含む15のモデルを評価する。全体精度が3%の下限を上回る12のモデルでは、対応する英語-韓国語の精度は±10ポイントの範囲内で統計的に同等であり(TOST)、提示言語だけによるコストはほとんどないことを示唆している。文字体系集約的なタスクではより明確な格差が見られる。韓国語暗号は英語より最大68.7ポイント低い一方、同じ字母を用いた暗号算は系統的なペナルティを示さず、字母合成の精度は韓国語暗号の精度を予測する。これらの対比は因果的というより診断的であり、多段階の音節以下の実行における困難さと整合する。韓国語のみのタスクは、符号が様々な規則適用の欠陥と、12モデル全てで正の親族関係の欠陥とを分離する。最後に、顕著なサイズ指標は15タイプ中7タイプでEasyからHardへ増加せず、構造的なサイズが経験的難易度の信頼できない代理指標となっていることを示す。
English
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.