ChatPaper.aiChatPaper

EduPanel: 教育用動画のための三エージェントLLM判定器 -- 信頼性、補完性、そして人間の信頼較正

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

July 20, 2026
著者: Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee
cs.AI

要旨

教育ビデオは主要な教育媒体となりつつあり、その教育的品質をスケーラブルに評価する必要性が高まっている。既存の自動評価器は、教育の質がマルチモーダルな証拠に依存し、普遍的な特性としてではなく、対象とする学習者に照らして評価されるべきであるという点で、この状況に完全には対応していない。我々はEduPanelを提案する。これはルーブリックに基づき、学習者条件付きで動作するLLM評価器であり、専門エージェント間で評価を分解し、教育品質の異なる側面について解釈可能な評価を生成する。専門家による研究、アーキテクチャのアブレーション実験、学習者ペルソナ分析を通じて、EduPanelは専門家中央値に匹敵する信頼性を達成した。専門家による評価では、そのフィードバックにより採点精度が向上し(MAE 0.87→0.73)、同時に専門家は信頼性の低い出力を盲目的に受け入れるのではなく検出できる(AUC = 0.77)。これらの結果は、EduPanelが人間専門家の代替ではなく、教育的評価のための効果的なアシスタントとして機能し得ることを示唆している。
English
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.