instruction-tuning knowledge-graph visual-grounding curriculum-aligned textual-supervision gaokao-mm curriculum-cognition education-llm
Abstract
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school. It contains nine node types and fourteen relation types covering curriculum structure and visual grounding. From this graph, we derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. We also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs. On K12-Bench, Gemini-3-Flash achieves only 57 percent exact match and Gemma-4-31B-IT reaches 46 percent, with Prereq and Neighbor being the hardest tasks. Our training experiments show that domain-specific supervision can reduce this gap. Under a matched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaokaoBench and EduEval. For vision-language models, K12-Train-Full achieves the best overall results on Gaokao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full DataFlow and WizardLM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary. We release the graph, benchmark, training data, and complete construction pipeline.
한국어 요약
한 줄 요약
K12-KGraph는 중국 교과서 기반의 과정 지식 그래프로, 교육용 LLM의 커리큘럼 인지 능력을 평가하고 훈련하는 데 활용된다.
핵심 기여도
- K12-KGraph: 중국 교과서에서 추출한 7개 노드 타입, 9개 관계 타입을 가진 과정 지식 그래프.
- K12-Bench: 23,640개 질문의 다중 선택 평가 벤치마크, 5가지 과제(Prereq, Neighbor 등)로 구성.
- K12-Train: 2,300개 샘플의 지도 학습 데이터, 텍스트 및 멀티모달 QA 쌍 포함.
- 실험 결과: K12-Train은 GaokaoBench, EduEval에서 기존 데이터셋보다 높은 성능을 보임.
핵심 아이디어
기존 교육용 LLM 평가가 단순 문제 풀이에 집중하는 반면, K12-KGraph는 커리큘럼 인지(curriculum cognition)라는 새로운 개념을 도입한다. 이는 지식의 구조적 이해, 예비 조건, 개념 분류, 실험-개념 연결, 교육적 순서 등을 포함한다. K12-KGraph는 중국 교과서에서 직접 추출된 지식 그래프로, 7개 노드 타입(Concept, Skill, Experiment 등)과 9개 관계 타입(prerequisites_for, relates_to 등)을 포함한다. 이를 기반으로 K12-Bench와 K12-Train을 생성하여, LLM이 커리큘럼 구조를 이해하고 학습할 수 있도록 유도한다.
기술적 접근법
- **K12-KGraph**: 중국 교과서(PEP)에서 추출한 지식 그래프.
- **노드 타입**: Concept, Skill, Experiment, Exercise, Section, Chapter, Book.
- **관계 타입**: prerequisites_for, relates_to, verifies, appears_in, leads_to 등.
- **K12-Bench**: 5가지 과제(Ground, Prereq, Neighbor, Evidence, Locate)로 구성된 23,640개 질문.
- **K12-Train**: 2,267개 텍스트 QA, 5,068개 멀티모달 VQA 샘플.
- **훈련 모델**: Qwen3-4B-Base, Llama3.1-8B-Base 사용.
- **하이퍼파라미터**: 2,300개 샘플 예산 내에서 훈련.
주요 결과
- **K12-Bench 성능**: Gemini-3-Flash 57%, Gemma-4-31B-IT 46% 정확도. 가장 어려운 과제는 Prereq와 Neighbor.
- **K12-Train 효과**: 2,300 샘플로 GaokaoBench, EduEval에서 기존 8개 데이터셋을 상회.
- **멀티모달 성능**: K12-Train-Full은 Gaokao-MM, MDK12-medium, K12Vista에서 최고 성능.
- **샘플 효율성**: DataFlow, WizardLM 기반 훈련보다 적은 샘플로 더 높은 성능.
의의 및 한계
K12-KGraph는 교육용 LLM이 단순 문제 풀이를 넘어 커리큘럼 구조를 이해하도록 유도하는 데 기여한다. 특히, K12-Train은 샘플 효율성이 높아, 제한된 데이터로도 성능 향상을 가능하게 한다. 그러나 K12-KGraph는 중국 교과서에만 기반하기 때문에, 다른 교육 체계에의 확장성은 명시되지 않음. 또한, 평가 과제는 다중 선택 중심으로, 실제 교육 상황에 대한 심층적 평가가 부족할 수 있다.
실용적 활용
K12-KGraph는 교육용 AI 개발, 토이트 트레이닝, 커리큘럼 기반 학습 시스템에 활용 가능하다. 특히, K12-Train은 훈련 데이터가 제한된 상황에서 효과적인 훈련 자료로 사용될 수 있다.