EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee

arXiv:2607.18529 · 2026-07-22 공개 · arXiv · PDF

multi-agent llm-judge teaching-videos rubric-grounded educational-evaluation learner-conditioned human-expert-evaluation pedagogical-quality

Abstract

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.

한국어 요약

한 줄 요약

EduPanel은 학습자 조건에 따른 교육 영상 평가를 위한 3개 에이전트로 구성된 LLM 평가 시스템으로, 전문가 수준의 신뢰도와 보완성을 보인다.

핵심 기여도

핵심 아이디어

EduPanel은 기존의 단일 LLM 평가 시스템과 달리, 교육 영상 평가를 3개의 전문 에이전트로 분할하여 각각의 평가 영역에 집중하도록 설계되었다. 이는 교육 영상이 다중 모달 정보(음성, 시각적 요소, 학습 목표 등)를 포함하고, 평가가 특정 학습자에 조건된다는 점을 반영한 것이다. 예를 들어, 수학적 설명은 고급 학습자에게는 유용하지만, 초보자에게는 부적절할 수 있다. EduPanel은 이러한 학습자 조건을 평가에 반영함으로써 보다 현실적인 평가를 가능하게 한다.

또한, EduPanel은 단순히 전문가와 동일한 결과를 내는 것을 넘어, 전문가가 놓친 문제점을 보완적으로 제시함으로써 평가의 질을 높인다. 이는 전문가가 AI의 피드백을 비판적으로 검토하면서 유용한 제안을 채택하고, 오류를 감지할 수 있도록 돕는다.

기술적 접근법

EduPanel은 다음과 같은 3개의 전문 에이전트로 구성된다:

EduPanel은 rubric-grounded 방식으로 평가하며, 각 평가 항목에 대해 **content map**, **발견된 문제**, **이유**를 함께 제공하여 평가의 해석 가능성을 높인다. 평가 시스템은 **multimodal video-language 모델**을 기반으로 하며, **학습자 프로필**을 조건으로 입력받아 평가를 수행한다.

주요 결과

의의 및 한계

EduPanel은 교육 영상 평가의 **확장성**과 **해석 가능성**을 동시에 달성한 첫 사례로, 전문가의 평가를 보완하는 **효과적인 보조 도구**로 활용 가능하다. 특히, 학습자 조건을 반영한 평가는 기존의 단일 평가 체계를 넘어서는 중요한 발전이다.

그러나 EduPanel은 **시각적 추론**이 필요한 평가 항목에서는 한계가 있으며, 영상의 시각적 요소를 완벽히 해석하지 못하는 경우가 있다. 또한, 평가 결과가 전문가의 판단을 완전히 대체하지 못하고, **비판적 검토가 필수적**이라는 점도 한계로 지적된다.

실용적 활용

EduPanel은 온라인 교육 콘텐츠 제작 시 **전문가 리뷰 비용을 절감**하고, **개별 학습자 맞춤 평가**를 가능하게 하는 도구로 활용될 수 있다. 특히, AI 생성 교육 콘텐츠의 품질 검증, 대규모 교육 플랫폼의 콘텐츠 선별, 교사의 콘텐츠 개선 지원 등에 유용하게 사용될 수 있다.