ChatPaper.aiChatPaper

NOLLI: 영어-한국어 성능 격차 진단을 위한 난이도 보정 퍼즐 벤치마크

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

August 5, 2026
저자: Dasol Choi, Joonyong Park, Daegon Yu, Soo Yong Kim, Youngsook Song, Seunghyeok Hong
cs.AI

초록

본 논문에서는 한국어 성능 격차가 발생하는 지점을 진단하기 위해 설계된 절차 생성(procedurally generated) 영어-한국어 퍼즐 벤치마크인 NOLLI를 소개한다. NOLLI는 15개 퍼즐 유형(25개 태스크, 7,500개 항목)으로 구성되며, 모든 인스턴스는 시드(seed)로 재생성 가능하고, 고유한 해답을 가짐이 검증되었으며, 결정적으로 채점된다. 난이도를 크기와 동일시하는 대신, 행동적으로 난이도를 보정하여 고정된 참조 모델이 목표 정확도 구간에 도달할 때까지 각 생성기를 조정한다. 3단계 설계는 일치하는 직접 번역, 한글 자모(음절 이하 단위의 문자)에 대한 문자 체계 적응, 그리고 한국어 문화 또는 표기법에 기반한 한국어 전용 태스크를 결합한다. 본 논문에서는 프런티어, 오픈 가중치, 한국 개발 모델 15개를 평가한다. 전체 정확도 하한 3% 이상인 12개 모델에서, 일치된 영어-한국어 정확도는 ±10 pp 범위 내에서 통계적으로 동등하며(TOST), 이는 제시 언어만으로 인한 비용이 거의 없음을 시사한다. 문자 체계 집약적 태스크에서는 더 뚜렷한 격차가 나타난다. 한국어 암호(Korean Cipher)는 영어 대비 최대 68.7 pp 낮은 반면, 동일한 자모를 사용하는 크립트산술(Cryptarithmetic)에서는 체계적 불이익이 나타나지 않았다. 또한 자모 조합(Jamo Composition) 정확도는 한국어 암호 정확도를 예측한다. 이러한 대비는 인과적이라기보다 진단적이며, 다단계 음절 이하 단위 실행의 어려움과 일치한다. 한국어 전용 태스크는 부호가 다양한 규칙 적용 결손과, 12개 모델 모두에서 양의 값을 보이는 친족 관계 결손을 구분한다. 마지막으로, 가시적 크기 척도는 15개 유형 중 7개에서 Easy에서 Hard로 갈수록 증가하지 않아, 구조적 크기가 경험적 난이도의 신뢰할 수 있는 대리 지표가 아님을 보여준다.
English
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.