Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng, Bin Liang, Huayu Deng, Yao Hu, Kam-Fai Wong, Mu Chuan

arXiv:2609.02998 · 2026-09-08 공개 · arXiv · PDF

code-generation on-policy-distillation grpo instruction-following mathematics gpu-utilization teacher-gating prompt-level

Abstract

On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.

한국어 요약

한 줄 요약

TGOPD는 토큰 수준의 지도 신뢰성을 프롬프트 단위로 검증하여 Vanilla OPD를 개선한 새로운 온-정책 디스틸레이션 방법이다.

핵심 기여도

핵심 아이디어

Vanilla OPD는 강한 teacher 모델이 학생의 rollout에 대해 토큰 수준의 reverse KL 신호를 제공함으로써 학습을 가속화하지만, teacher의 신뢰도를 프롬프트 단위로 검증하지 않는다. 이는 teacher가 확신을 갖고 있지만 오류가 있는 경우 학생이 잘못된 신호를 학습하게 되는 문제를 유발한다. TGOPD는 teacher가 idle 상태일 때 **K<sub>T</sub>개의 probe rollout**을 생성하고, **verifier**가 이를 평가하여 **q<sub>T</sub>(x) = K<sub>T</sub><sup>-1</sup>Σ<sub>k</sub>r<sub>k</sub>**로 신뢰도를 추정한다. 이 신뢰도가 임계치 τ 이상이면 **dense OPD**를 적용하고, 그렇지 않으면 **verifier-grounded GRPO**로 라우팅함으로써 teacher 신호를 선택적으로 제어한다. 이는 기존의 distributional proxy(예: entropy, agreement)가 아닌 **outcome-based reliability estimation**을 도입한 핵심 차별점이다.

기술적 접근법

주요 결과

의의 및 한계

TGOPD는 teacher의 idle capacity를 활용하여 신뢰도 평가를 수행함으로써 학습 신호의 질과 계산 효율성을 동시에 개선한 점에서 학술적·실용적 가치가 있다. 기존의 distributional proxy(예: entropy, agreement)는 teacher의 정확성 자체를 평가하지 못하는 한계를 극복하고, outcome-based 신뢰도 평가를 도입한 점이 혁신적이다. 그러나 현재의 구현은 **binary routing**만 지원하며, **open-ended task**나 **uncertainty-aware gate**로의 확장은 여전히 개선이 필요한 부분이다. 또한, **verifier 의존성**이 높아 verifier의 품질이 최종 성능에 큰 영향을 미친다.

실용적 활용

TGOPD는 대규모 언어 모델의 post-training 과정에서 teacher 신호의 신뢰도를 프롬프트 단위로 검증하는 데 유용하며, 특히 **코드 생성**, **수학 문제 풀이**, **instruction following**과 같은 정확도가 중요한 도메인에서 활용 가능하다. 계산 자원 활용률 향상 덕분에 **GPU 클러스터 운영 비용**을 줄이는 데도 기여할 수 있다.