NOLLI:用于诊断英韩性能差距的难度校准谜题基准
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
August 5, 2026
作者: Dasol Choi, Joonyong Park, Daegon Yu, Soo Yong Kim, Youngsook Song, Seunghyeok Hong
cs.AI
摘要
我们提出 NOLLI,一个程序化生成的英韩谜题基准,旨在诊断韩语性能差距的来源。该基准包含 15 种谜题类型(25 个任务;7,500 个条目),每个实例均可通过种子重新生成、经验证具有唯一解,并以确定性方式评分。我们并不将“更难”等同于“更大”,而是通过行为表现来标定难度,调整每个生成器,直到固定参考模型落入目标准确率区间。其三层设计结合了匹配的直接翻译、基于韩文字母(音节下字母)的文字系统改编,以及植根于韩国文化或正字法的仅韩语任务。我们评估了 15 个前沿模型、开放权重模型和韩国开发的模型;在总体准确率下限超过 3% 的 12 个模型中,匹配的英韩准确率在 ±10 个百分点范围内统计等价(TOST),表明仅呈现语言本身带来的代价很小。高度依赖书写系统的任务展现出更明显的差距:韩语密码任务落后英语最多 68.7 个百分点,而使用相同韩文字母的字母算式任务则没有系统性劣势;韩文字母组合准确率可以预测韩语密码任务准确率。这些对比具有诊断意义而非因果意义,与多步骤亚音节执行的难度相一致。仅韩语任务将规则应用缺陷(方向各异)与亲属关系缺陷(在所有 12 个模型中均为正值)区分开来。最后,一个显著的规模度量在 15 种类型中有 7 种未能随难度从易到难增长,表明结构性规模难以作为经验难度的可靠代理指标。
English
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.