ChatPaper.aiChatPaper

EduPanel:用於教學影片的三代理LLM評判者——可靠性、互補性與人類信任校準

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

July 20, 2026
作者: Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee
cs.AI

摘要

教學影片正逐漸成為主要的教育媒介,因此對於其教學品質進行可擴展評估的需求也日益增長。現有的自動評判系統並未能完全應對此情境,因為教學品質取決於多模態證據,且應根據預期的學習者來評估,而非將其視為普遍屬性。我們提出了EduPanel,這是一個基於評分量規、以學習者為條件的LLM評判系統,它將評估任務分解到專門的代理中,從而為教學品質的不同層面提供可解釋的評估結果。透過專家研究、架構消融實驗以及學習者角色分析,EduPanel達到了與人類專家中位數相當的信度。在專家評估中,其反饋提升了評分準確度(平均絕對誤差從0.87降至0.73),同時專家仍能偵測到不可靠的輸出(AUC = 0.77),而非盲目接受。這些結果表明,EduPanel可作為教育評估的有效輔助工具,而非取代人類專家。
English
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.