EduPanel: 세 에이전트 LLM 평가자를 활용한 교육 동영상 평가 -- 신뢰성, 상보성 및 인간 신뢰도 보정
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
July 20, 2026
저자: Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee
cs.AI
초록
교육 비디오가 교육의 주요 매체로 자리 잡으면서, 이들의 교육학적 품질에 대한 확장 가능한 평가의 필요성이 증가하고 있다. 기존의 자동 평가 시스템은 이러한 상황을 완전히 해결하지 못하는데, 이는 교수 품질이 다중 양식 증거에 의존하며, 보편적 속성이 아닌 대상 학습자(Frms)를 기준으로 평가되어야 하기 때문이다. 본 논문에서는 루브릭 기반 학습자 조건부 LLM 평가 시스템인 EduPanel을 제안한다. 이 시스템은 평가를 전문화된 에이전트들로 분해하여 교수 품질의 다양한 측면에 대해 해석 가능한 평가를 산출한다. 전문가 연구, 아키텍처 절제 실험 및 학습자 페르소나 분석을 통해 EduPanel은 중간 인간 전문가에 필적하는 신뢰성을 달성했다. 전문가 평가에서 EduPanel의 피드백은 점수 정확도를 향상시켰으며(MAE 0.87 → 0.73), 전문가들은 신뢰할 수 없는 출력을 맹목적으로 수용하지 않고 탐지할 수 있었다(AUC = 0.77). 이러한 결과는 EduPanel이 인간 전문가를 대체하는 것이 아니라 교육 평가를 위한 효과적인 보조 도구로 기능할 수 있음을 시사한다.
English
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.