K12-KGraph:用於教育大型語言模型基準測試與訓練的課程對齊知識圖譜
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
July 23, 2026
作者: Hao Liang, Qihan Lin, Zhaoyang Han, Xiaochen Ma, Zhen Hao Wong, Meiyi Qiang, Linzhuang Sun, Wentao Zhang
cs.AI
摘要
大型語言模型日益廣泛應用於K-12教育,然而現有基準主要測試應試問答能力,而非理解課程知識如何被結構化與視覺化呈現。我們將此能力稱為「課程認知」,涵蓋先決條件鏈、概念分類體系、實驗-概念連結、教學排序及視覺基礎。我們提出K12-KGraph,這是一個根據官方人民教育出版社中小學數學、物理、化學與生物課本所擷取、對齊課程的知識圖譜,包含九種節點類型與十四種關係類型,涵蓋課程結構與視覺基礎。由此圖譜衍生出K12-Bench,一個包含23,640道多選題的基準測試,涵蓋五大任務族:基礎定位(Ground)、先決條件(Prereq)、鄰近概念(Neighbor)、證據(Evidence)與定位(Locate)。我們亦建構K12-Train,一個由7,335個樣本組成的圖譜引導監督微調語料庫,其中包含2,267個純文字問答對與5,068個多模態視覺問答對。在K12-Bench上,Gemini-3-Flash僅達57%完全匹配,Gemma-4-31B-IT則達46%,其中先決條件與鄰近概念為最困難任務。我們的訓練實驗顯示,領域特定的監督資料可縮小此差距。在匹配的2,300樣本預算下,K12-Train-Text在GaokaoBench與EduEval上始終優於八個主流指令微調語料庫中同等大小的子集。對於視覺語言模型,儘管K12-Train-Full使用的樣本數少於完整的DataFlow與WizardLM基準線,仍在Gaokao-MM、MDK12-medium與K12Vista上取得所有比較訓練配置中最佳的整體結果。同時,其表現超越純文字與純多模態變體,顯示文字與視覺監督具有互補性。我們公開釋出該圖譜、基準測試、訓練資料及完整的建構流程。
English
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school. It contains nine node types and fourteen relation types covering curriculum structure and visual grounding. From this graph, we derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. We also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs. On K12-Bench, Gemini-3-Flash achieves only 57 percent exact match and Gemma-4-31B-IT reaches 46 percent, with Prereq and Neighbor being the hardest tasks. Our training experiments show that domain-specific supervision can reduce this gap. Under a matched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaokaoBench and EduEval. For vision-language models, K12-Train-Full achieves the best overall results on Gaokao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full DataFlow and WizardLM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary. We release the graph, benchmark, training data, and complete construction pipeline.