DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

Karn Tiwari, Varnith Chordia, Prathosh A P

arXiv:2610.04596 · 2026-10-07 공개 · arXiv · PDF

reinforcement-learning large-language-models code-generation on-policy-distillation group-relative-policy-optimization trajectory-level-supervision solution-coverage diffgate

Abstract

On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but suffers from sparse rewards and coarse credit assignment. We show that OPD and RLVR exhibit complementary blind spots: teacher signals provide dense local guidance but are weakly aligned with rollout correctness, whereas group-relative rewards capture task success but provide coarse token-level credit and vanish on all-failure groups. We introduce DiffGate, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance. Teacher supervision is applied only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher--student discrepancies from dominating optimization. The verifier therefore determines which trajectories receive teacher guidance, while the teacher provides dense token-level update directions within those trajectories. Across Qwen3-0.6B and Qwen3-1.7B students, DiffGate improves code avg@8 over matched GRPO by +1.7 and +1.8 points and pass@8 by +1.6 and +5.7 points, respectively. On mathematics, avg@8 remains within 0.5 points of GRPO while pass@8 improves by +1.1 and +3.9 points. Overall, DiffGate improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under our evaluation protocol.

한국어 요약

한 줄 요약

DiffGate는 GRPO와 선택적, 제한된 교사 지도를 결합한 새로운 온-포리시 디스틸레이션 방법으로, 실패한 추적에만 교사 신호를 적용하여 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

기존의 온-포리시 디스틸레이션(OPD)은 토큰 수준의 밀집된 지도를 제공하지만, 추적 전체의 정확도와는 약하게 연관된다. 반면, GRPO는 추적 수준의 성공 여부를 반영하지만 토큰 수준의 세부 신호가 부족하고, 모든 실패 그룹에서는 학습 신호가 사라진다. 이 두 접근법의 보완적인 단점을 해결하기 위해, DiffGate는 결과 기반으로 교사 지도를 선택적으로 적용하는 새로운 학습 목적 함수를 제안한다.

DiffGate는 GRPO와 선택적, 제한된 교사 지도를 결합하여, 실패한 추적에만 교사 지도를 적용하고, 그룹 난이도에 따라 스케일링한다. 이는 GRPO가 신호를 제공하지 않는 모든 실패 그룹에서 특히 유용하며, 교사와 학생 간의 큰 차이가 업데이트를 지배하지 않도록 보장한다. 이 방식은 교사가 밀집된 토큰 수준의 업데이트 방향을 제공하면서도, 결과 피드백이 약한 영역에 집중적으로 지도를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

DiffGate는 기존 OPD와 GRPO의 단점을 보완하여, 추적 수준의 성공 여부와 토큰 수준의 밀집된 지도를 동시에 고려하는 새로운 학습 목적 함수를 제안한다. 이는 학습 신호가 약한 영역에 집중적으로 지도를 제공함으로써, 학생 모델의 해결 범위(solution coverage)를 향상시킨다. 특히, 모든 실패 그룹에서 GRPO가 신호를 제공하지 않는 경우에도 교사 지도를 적용할 수 있어, 학습 효과를 극대화한다.

그러나, DiffGate는 여전히 교사 모델의 품질에 의존하며, 교사와 학생 간의 큰 차이가 업데이트에 영향을 줄 수 있다는 한계가 있다. 또한, 그룹 크기와 난이도 조절이 학습 효과에 중요한 역할을 하므로, 이러한 하이퍼파라미터의 최적화가 필요하다.

실용적 활용

DiffGate는 코드 생성, 수학 문제 해결 등 추론이 필요한 자연어 생성 작업에 적용 가능하다. 특히, 학습 신호가 약한 영역에서 성능 향상이 필요한 경우, GRPO와 OPD를 결합한 이 접근법은 효과적인 학습 전략이 될 수 있다. 또한, 대규모 언어 모델의 후처리 훈련(post-training)에서 활용하여, 학습-추론 불일치 문제를 완화할 수 있다.