ChatPaper.aiChatPaper

K12-KGraph: カリキュラム準拠の知識グラフ - 教育用LLMのベンチマークと訓練のための

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

July 23, 2026
著者: Hao Liang, Qihan Lin, Zhaoyang Han, Xiaochen Ma, Zhen Hao Wong, Meiyi Qiang, Linzhuang Sun, Wentao Zhang
cs.AI

要旨

大規模言語モデルはK-12教育においてますます活用されているが、既存のベンチマークは主に試験問題の解答能力を評価するものであり、カリキュラム知識がどのように構造化され視覚的に提示されるかを理解する能力は評価していない。我々はこの能力を「カリキュラム認知」と呼ぶ。これには、前提条件の連鎖、概念の分類体系、実験と概念の関連、教育的な順序付け、視覚的基盤付けが含まれる。我々はK12-KGraphを導入する。これは、小学校、中学校、高校向けの数学、物理、化学、生物学の公式な人民教育出版社の教科書から抽出された、カリキュラムに沿った知識グラフである。このグラフは9種類のノードタイプと14種類のリレーションタイプを含み、カリキュラム構造と視覚的基盤付けを網羅している。このグラフから、我々はK12-Benchを導出する。これは23,640問の多肢選択式ベンチマークであり、Ground、Prereq、Neighbor、Evidence、Locateの5つのタスクファミリーを含む。さらに、7,335サンプルからなるグラフ誘導型の教師ありファインチューニングコーパスK12-Trainを構築する。このコーパスには2,267件のテキストのみのQAペアと5,068件のマルチモーダルVQAペアが含まれる。K12-Benchにおいて、Gemini-3-Flashは57%の完全一致率しか達成できず、Gemma-4-31B-ITは46%に留まり、PrereqとNeighborが最も難しいタスクである。我々の訓練実験は、ドメイン固有の教師信号がこのギャップを縮小できることを示している。2,300サンプルという同等の予算の下で、K12-Train-Textは、GaokaoBenchとEduEvalにおいて、8つの主流の指示チューニングコーパスの同等サイズのサブセットを一貫して上回る。視覚言語モデルに関しては、K12-Train-Fullは、比較したすべての訓練構成の中で、Gaokao-MM、MDK12-medium、K12Vistaにおいて最良の総合結果を達成している。これは、完全なDataFlowおよびWizardLMベースラインよりも少ないサンプル数であるにもかかわらずである。また、テキストのみ、マルチモーダルのみのバリアントをも上回っており、テキストと視覚の教師信号が相補的であることを示している。我々は、グラフ、ベンチマーク、訓練データ、および完全な構築パイプラインを公開する。
English
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school. It contains nine node types and fourteen relation types covering curriculum structure and visual grounding. From this graph, we derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. We also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs. On K12-Bench, Gemini-3-Flash achieves only 57 percent exact match and Gemma-4-31B-IT reaches 46 percent, with Prereq and Neighbor being the hardest tasks. Our training experiments show that domain-specific supervision can reduce this gap. Under a matched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaokaoBench and EduEval. For vision-language models, K12-Train-Full achieves the best overall results on Gaokao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full DataFlow and WizardLM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary. We release the graph, benchmark, training data, and complete construction pipeline.