ChatPaper.aiChatPaper

NOLLI:一個難度校準的謎題基準,用於診斷英語-韓語效能差距

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

August 5, 2026
作者: Dasol Choi, Joonyong Park, Daegon Yu, Soo Yong Kim, Youngsook Song, Seunghyeok Hong
cs.AI

摘要

我們介紹 NOLLI——一個程序化生成的英韓謎題基準,旨在診斷韓文效能差距的來源。它包含 15 種謎題類型(25 項任務;7,500 個題項),每個實例皆可透過種子重新生成、經驗證具有唯一解,並以確定性方式計分。我們並未將「更難」等同於「更大」,而是以行為層面校準難度,調整每個生成器,直到固定參考模型落入目標準確度區間。其三層設計結合了配對的直接翻譯、基於韓文 jamo(亞音節字母)的文字改編,以及根植於韓國文化或正字法的韓文專屬任務。 我們評估了 15 個前沿、開放權重及韓國開發的模型;在 12 個高於 3% 總體準確度下限的模型中,配對英韓準確度在正負 10 個百分點範圍內統計上等效(TOST),顯示單就呈現語言而言幾乎沒有代價。書寫系統密集度高的任務顯示出更明顯的差距:韓文密碼(Korean Cipher)落後英文最多達 68.7 個百分點,而使用相同 jamo 的密碼算術(Cryptarithmetic)卻沒有系統性的劣勢;此外,Jamo 組合的準確度可預測韓文密碼的準確度。這些對比是診斷性而非因果性的,與多步驟亞音節執行的困難度一致。韓文專屬任務將規則應用缺陷(其方向不一)與在所有 12 個模型中皆呈正向的親屬稱謂缺陷區分開來。最後,一個顯著的規模量度在 15 種類型中有 7 種未能從「簡單」到「困難」隨之增長,使結構規模成為經驗難度的不可靠代理指標。
English
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.