D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang

arXiv:2608.24987 · 2026-08-27 공개 · arXiv · PDF

on-policy convergence-rate multi-teacher-distillation reverse-kl-divergence dynamic-domain-scheduling qwen3-6-35b-a3b domain-expert-teachers rollout-steps

Abstract

Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D$^3$-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D$^3$-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D$^3$-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3$\times$ reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.

한국어 요약

한 줄 요약

D$^3$-MOPD는 학습 속도에 따라 도메인 혼합을 동적으로 조정해 MOPD의 효율성을 3× 향상시키는 제로오버헤드 스케줄러이다.

핵심 기여도

핵심 아이디어

기존 MOPD는 도메인별 학습 속도 차이를 무시하고 고정된 도메인 혼합을 사용해, 일부 도메인은 과소학습되고 다른 도메인은 과잉 학습된다. 이에 D$^3$-MOPD는 학습 중 생성되는 reverse-KL 신호를 활용해 도메인별 학습 진행도를 추적하고, 이를 바탕으로 도메인 샘플링 비율을 실시간으로 조정한다. 핵심 아이디어는 학습 중 발생하는 KL 발산을 "학습 상태 지표"로 재사용하는 것이다. 이는 핵심 학습 루프를 수정하지 않으면서도, 외부 watcher 프로세스를 통해 도메인별 KL 히스토리를 분석하고, KL 갭과 최근 감소 속도를 결합한 composite signal을 생성해 softmax-floor 함수를 통해 샘플링 비율을 업데이트한다.

기술적 접근법

주요 결과

의의 및 한계

D$^3$-MOPD는 MOPD의 핵심 문제인 도메인별 학습 속도 불균형을 해결함으로써, 학습 효율과 최종 모델 성능을 동시에 향상시킨다. 기존 정적 혼합 방식 대비 3× 빠른 학습 속도와 97%의 성능 격차 축소는 실용적 가치가 높다. 또한, 핵심 학습 루프를 변경하지 않아 기존 MOPD 연구와의 호환성이 높으며, 다양한 loss function 및 교사 학습 기법과 결합 가능하다는 장점이 있다. 다만, D$^3$-MOPD는 reverse-KL 신호에만 의존하므로, KL이 학습 상태를 정확히 반영하지 못하는 경우 성능이 저하될 수 있다. 또한, watcher의 주기적 업데이트가 일정 지연을 유발할 수 있으나, 본 연구에서는 이를 최소화한 설계를 제시하고 있다.

실용적 활용

D$^3$-MOPD는 다중 도메인에서 전문성 있는 대형 언어 모델을 효율적으로 학습시키는 데 활용 가능하다. 특히, 수학, 코드 생성, 인스트럭션 팔로잉, 툴 사용 등 다양한 도메인을 아우르는 모델 개발에 적합하며, 학습 시간과 컴퓨팅 자원을 절감할 수 있어 산업 현장에서의 모델 트레이닝 효율화에 기여할 수 있다.