SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

Haonan He, Haodi Lei, Yun Luo, Haoran Zhang, Shunkai Zhang, Yizhuo Li, Shengji Tang, Zhilin Wang, Runzhe Zhan, Lei Bai, Ganqu Cui, Fangchen Yu, Yafu Li, Peng Ye, Ning Ding, Yu Cheng

arXiv:2608.14277 · 2026-08-17 공개 · arXiv · PDF

on-policy-distillation mathematical-reasoning qwen3 science-benchmarks proof-reasoning long-context-reasoning tokenizer-agnostic student-reference-kl-loss

Abstract

On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mismatch, teacher-student distribution mismatch, response length explosion, and training instability. In this work, we study this setting by transferring proof-reasoning capabilities from the long-context reasoning model SU-01 to short-context student models. To handle tokenizer differences, we perform OPD in a shared text space and align only tokens that occupy identical text spans under the student and teacher tokenizers. To mitigate the problem of excessive generation length and frequent truncation, we introduce a student reference KL loss and mask the advantages of special termination tokens such as </think> and <|im_end|>. This strategy constrains the student from drifting excessively from its initial policy, thereby mitigating the teacher-student distribution mismatch problem and fostering steady length growth. Experiments on both same-family and different-family student models, including Qwen3, Qwen3.5, Intern-S2, GLM-4.7, Gemma-4, show consistent gains in mathematical reasoning, especially natural-language math proving. Notably, Intern-S2-Preview improves by 21.2 points on ProofBench, reaching 55.2 and surpassing Gemini-2.5-Pro. It also improves on science benchmarks such as HLE and HiPhO, suggesting that OPD transfers reasoning capabilities that generalize beyond the mathematical training domain.

한국어 요약

한 줄 요약

SimpleOPD는 토크나이저 차이를 고려한 온-폴리시 디스틸레이션으로 수학 증명 능력을 전이시켜 모델 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

SimpleOPD는 강력한 토크나이저-비종속 온-폴리시 디스틸레이션 기법으로, 수학적 추론 능력을 전이시킨다. 기존 온-폴리시 디스틸레이션(OPD)은 교사 모델과 학생 모델 간 토크나이저 불일치, 분포 불일치, 길이 폭발, 훈련 불안정성 등의 문제를 해결하지 못했다. 본 연구는 SU-01이라는 긴 컨텍스트 추론 모델을 기반으로, 학생 모델이 짧은 컨텍스트를 사용하는 경우에도 효과적인 디스틸레이션이 가능하도록 설계했다.

핵심 아이디어는 공유 텍스트 공간에서 토큰 정렬을 수행하는 것이다. 즉, 학생과 교사 토크나이저가 동일한 텍스트 스팬을 생성하는 토큰만 정렬하여, 토크나이저 차이를 무시하면서도 신뢰할 수 있는 토큰 수준의 지도를 제공한다. 또한, 종결 토큰(`< /think >`, `< |im_end| >`)의 이점을 마스킹하고, 학생 정책의 KL 손실을 도입하여 학생이 초기 정책에서 과도하게 벗어나는 것을 방지함으로써 길이 폭발과 훈련 불안정성을 완화한다.

기술적 접근법

주요 결과

의의 및 한계

SimpleOPD는 토크나이저 차이가 있는 교사-학생 모델 간에도 수학적 추론 능력을 전이시킬 수 있음을 보여준다. 특히, 토크나이저-비종속 디스틸레이션을 통해 교사 모델의 긴 컨텍스트 추론 능력을 짧은 컨텍스트 학생 모델로 전이시킬 수 있으며, 이는 기존의 토크나이저 일치 조건 없이도 가능하다는 점에서 학술적 의의가 있다. 또한, 종결 토큰 마스킹과 KL 손실을 통해 훈련 안정성과 길이 제어를 동시에 달성한 점도 실용적으로 유용하다.

그러나 토크나이저 차이가 클수록 전이 효과가 감소하는 한계가 있다. 예를 들어, Gemma-4-26B-A4B는 SentencePiece 기반 토크나이저를 사용하여 SU-01과의 토크나이저 차이가 크고, 이로 인해 AnswerBench 성능이 오히려 감소했다. 따라서 토크나이저 일치 정도가 전이 성능에 중요한 영향을 미친다는 점이 드러난다.

실용적 활용

SimpleOPD는 수학 증명, 과학 문제 해결 등 복잡한 추론이 필요한 분야에서 모델 성능을 향상시키는 데 활용될 수 있다. 특히, 다양한 토크나이저와 아키텍처를 가진 모델 간의 추론