ChatPaper.aiChatPaper

CogEvol: 효율적이고 신뢰할 수 있는 학습 환경 생성을 향하여

CogEvol: Towards Efficient and Reliable Learning Environment Generation

August 31, 2026
저자: Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Haoxuan Li, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang
cs.AI

초록

우리는 학습 환경 생성을 위해 특별히 훈련된 모델군인 CogEvol을 제시한다. CogEvol은 강의 개요를 단일 패스로 완성된 학습 산출물(구조화된 JSON 슬라이드 또는 독립형 대화형 HTML 페이지)로 변환한다. 22만 건의 프로덕션 요청에서 CogEvol은 슬라이드를 중앙값 17초, 대화형 페이지를 59초 만에 완성하여 수 분이 소요되던 다중 턴 에이전트 스캐폴딩을 대체한다. 신뢰성은 기대가 아닌 강제로 확보된다. 프로덕션 기반 데이터 파이프라인은 실제 실패 사례를 53,687개의 검증된 SFT 샘플로 전환하며, 하이브리드 규칙-플러스-VLM 보상이 GRPO 기반 RL을 구동한다. 이 과정에서 시각적으로는 그럴듯하지만 플레이 불가능한 게임을 생성하는 보상 해킹 사례를 발견·수정한 후 강화가 이루어졌다. CogEvol-27B는 슬라이드 품질 83.7점, 500건의 대화형 HTML 벤치마크에서 63.7점을 기록하며, 주력 코딩 모델 대비 26.9배 적은 파라미터로 동일한 성능을 달성한다. 또한 OpenMAIC 팀과의 협업을 통해 해당 팀의 실시간 프로덕션 트래픽을 처리한다. CogEvol-4B는 Apache 2.0 라이선스로 https://github.com/CogEvol/CogEvol-4B 에서 공개되었으며, 외부 주력 모델들은 동일한 하니스로 동일한 평가 스위트에서 측정된다. 스캐폴드 편집은 대화형 페이지 생성 비용을 약 76% 추가 절감하며, 전체 스택은 국내산 Ascend 가속기에서 A800 GPU와 애플리케이션 수준의 동등성을 달성하여 AI 기반 교육의 대규모 단위 비용을 낮춘다.
English
We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.