Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie, Jinghao Guo, Xingwei Qu, Yichi Zhang, Chau Yuen

arXiv:2609.35347 · 2026-09-29 공개 · arXiv · PDF

reinforcement-learning benchmark-evaluation language-models on-policy-distillation instruction-following multi-teacher mathematics-specialist domain-normalization

Abstract

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.

한국어 요약

한 줄 요약

DN-MOPD는 전문가 모델들의 피드백 스케일을 정규화하여 MOPD의 성능 한계를 극복하고, 수학 능력 향상과 다중 태스크 성능 향상을 달성한다.

핵심 기여도

핵심 아이디어

MOPD는 도메인별 전문가가 학습자의 응답에 대해 토큰 단위 피드백을 제공하는 방식으로 전문가 모델을 통합하지만, 피드백의 스케일 차이를 고려하지 않아 학습 불균형이 발생한다. 예를 들어, 인스트럭션 팔로잉 전문가의 로그-비율 분산은 수학 전문가의 2.3~4.4배이며, 이는 학습자의 합성 그라디언트에 94%를 차지한다. 이는 단순히 어떤 전문가가 피드백을 제공하는지(라우팅)가 아니라, 그 피드백이 얼마나 강하게 반영되는지도 결정해야 한다는 점을 드러낸다. DN-MOPD는 도메인별 피드백의 로그-비율 분산을 배치별로 측정하고, 이를 기반으로 스케일을 조정하여 불균형을 해소한다. 이는 추가적인 모델이나 라우터 없이 MOPD에 단일 연산만 추가하는 방식으로, MOPD는 모든 스케일러가 1인 특별한 경우가 된다.

기술적 접근법

주요 결과

의의 및 한계

DN-MOPD는 전문가 모델 통합에서 피드백 스케일 불균형 문제를 해결함으로써, MOPD의 성능 한계를 극복하고 수학 및 다중 태스크 성능을 향상시킨다. 특히, 인스트럭션 팔로잉 전문가의 지나친 영향력을 억제함으로써 학습 불균형을 해소하는 점이 학술적·실용적 가치가 있다. 그러나 본 연구는 Qwen3.5 모델 내에서 한정된 전문가 풀을 사용했으며, 학습자 재훈련 없이 비교했기 때문에 일반화 가능성에 한계가 있다. 또한, 스케일 추정은 도메인 구성과 응답 길이에 따라 달라지고, 클리핑 경계는 크기별로 조정되지 않았다.

실용적 활용

DN-MOPD는 수학, 코드 생성, 인스트럭션 팔로잉 등 다양한 도메인에서 전문가 모델을 통합해야 하는 대형 언어 모델의 후처리 단계에 적용 가능하다. 특히, 인스트럭션 팔로잉 전문가의 지나친 영향력이 학습 불균형을 유발하는 경우, DN-MOPD는 모델 성능을 안정적으로 향상시키는 데 유용하다.