ChatPaper.aiChatPaper

K12-KGraph: 교육 과정에 맞춰진 지식 그래프로, 교육용 LLM의 벤치마킹 및 훈련을 위한 도구

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

July 23, 2026
저자: Hao Liang, Qihan Lin, Zhaoyang Han, Xiaochen Ma, Zhen Hao Wong, Meiyi Qiang, Linzhuang Sun, Wentao Zhang
cs.AI

초록

대규모 언어 모델은 초중등 교육(K-12)에서 점점 더 많이 사용되고 있지만, 기존 벤치마크는 주로 시험 문제 응답 능력을 평가할 뿐 교과 지식이 어떻게 구조화되고 시각적으로 제시되는지에 대한 이해는 평가하지 않습니다. 우리는 이러한 능력을 **커리큘럼 인지(curriculum cognition)** 라고 부릅니다. 이는 전제 조건 체인(prerequisite chains), 개념 분류 체계(concept taxonomies), 실험-개념 연결(experiment-concept links), 교수 순서(pedagogical sequencing), 그리고 시각적 접지(visual grounding)를 포괄합니다. 우리는 중국 인민교육출판사(PEP) 공식 교과서(초등, 중등, 고등학교 수학, 물리, 화학, 생물)에서 추출한 커리큘럼 정렬 지식 그래프인 **K12-KGraph**를 소개합니다. 이 그래프는 9개의 노드 유형과 커리큘럼 구조 및 시각적 접지를 포괄하는 14개의 관계 유형을 포함합니다. 이 그래프로부터 우리는 다섯 가지 작업군( Ground, Prereq, Neighbor, Evidence, Locate)으로 구성된 23,640개의 객관식 문항 벤치마크인 **K12-Bench**를 도출합니다. 또한 그래프 기반 지도 미세 조정을 위한 7,335개 샘플로 구성된 **K12-Train** 코퍼스를 구축했으며, 여기에는 2,267개의 텍스트 전용 QA 쌍과 5,068개의 멀티모달 VQA 쌍이 포함됩니다. K12-Bench에서 Gemini-3-Flash는 정확 일치(exact match) 57%에 불과했고, Gemma-4-31B-IT는 46%를 기록했으며, Prereq와 Neighbor 작업이 가장 어려운 것으로 나타났습니다. 우리의 훈련 실험은 도메인 특화 지도 학습이 이러한 성능 격차를 줄일 수 있음을 보여줍니다. 2,300개 샘플의 동일 예산 조건에서 K12-Train-Text는 GaokaoBench와 EduEval에서 8개의 주요 명령어 튜닝 코퍼스의 동일 크기 부분집합보다 일관되게 우수한 성능을 보였습니다. 시각-언어 모델의 경우, K12-Train-Full은 전체 DataFlow 및 WizardLM 베이스라인보다 적은 샘플을 사용했음에도 불구하고 모든 비교 훈련 구성 중 Gaokao-MM, MDK12-medium, K12Vista에서 최고의 전체 성능을 달성했습니다. 또한 텍스트 전용 및 멀티모달 전용 변형을 모두 능가하여 텍스트 및 시각적 지도 학습이 상호 보완적임을 보여줍니다. 우리는 그래프, 벤치마크, 훈련 데이터 및 완전한 구축 파이프라인을 공개합니다.
English
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school. It contains nine node types and fourteen relation types covering curriculum structure and visual grounding. From this graph, we derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. We also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs. On K12-Bench, Gemini-3-Flash achieves only 57 percent exact match and Gemma-4-31B-IT reaches 46 percent, with Prereq and Neighbor being the hardest tasks. Our training experiments show that domain-specific supervision can reduce this gap. Under a matched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaokaoBench and EduEval. For vision-language models, K12-Train-Full achieves the best overall results on Gaokao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full DataFlow and WizardLM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary. We release the graph, benchmark, training data, and complete construction pipeline.