CogEvol: Towards Efficient and Reliable Learning Environment Generation

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Haoxuan Li, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang

arXiv:2608.30968 · 2026-09-01 공개 · arXiv · PDF

grpo rl-training model-optimization education-ai sft-samples ascend-accelerators cogevol learning-environment-generation

Abstract

We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.

한국어 요약

한 줄 요약

CogEvol은 교육 환경 생성을 위한 단일 패스 모델로, 220k 요청에서 슬라이드 생성 시간을 17초로 단축하고 신뢰성을 강화한 RL 기반 시스템이다.

핵심 기여도

핵심 아이디어

기존 교육 자료 생성은 다중 턴 에이전트를 사용해 수분이 소요되지만, CogEvol은 단일 패스로 JSON 슬라이드나 HTML 페이지를 생성한다. 이는 학습 환경 생성의 **속도**와 **신뢰성** 문제를 해결하기 위한 핵심 전략이다. 또한, 실패 사례를 기반으로 53,687개의 SFT 샘플을 생성해 모델의 신뢰성을 강제적으로 보장한다. 보상 함수는 VLM과 규칙 기반 혼합으로 설계되며, GRPO 기반 강화 학습을 통해 보상 해킹을 방지한다. 특히, 시각적으로는 완성되어 보이지만 실행 불가능한 게임을 생성한 사례를 통해 보상 설계의 중요성을 강조한다.

기술적 접근법

주요 결과

의의 및 한계

CogEvol은 대규모 AI 교육 환경 생성의 **속도**, **신뢰성**, **비용** 문제를 동시에 해결하며, 실제 교육 현장에서 사용 가능한 솔루션을 제시한다. 특히, **보상 설계의 중요성**을 강조하며, 시각적 판단만으로는 측정 불가능한 실패를 방지하는 기술적 접근법을 제시한다. 그러나, **하이퍼파라미터 세부 사항**이나 **모델 아키텍처 구조**는 명시되지 않아 학술적 재현성에 한계가 있을 수 있다.

실용적 활용

CogEvol은 온라인 교육 플랫폼, 학교 교재 자동 생성, 개별 맞춤형 학습 환경 제작 등에서 활용 가능하다. 특히, 저비용으로 대규모 AI 교육을 구현하고자 하는 기관이나 개발자에게 유용하며, CogEvol-4B의 오픈소스 배포를 통해 접근성이 높아졌다.