Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li, Wenfeng Feng, Guohua Liu, Yuewei Zhang

arXiv:2610.08448 · 2026-10-07 공개 · arXiv · PDF

code-generation mathematical-reasoning on-policy-distillation reverse-kl gradient-analysis token-mismatch tokenizer-alignment span-log-probabilities

Abstract

On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.

한국어 요약

한 줄 요약

크로스-토크나이저 온-포로시 디스틸레이션에서 정렬 범위보다 감독 신뢰도가 학습 성능에 더 중요한 역할을 한다는 점을 실증적으로 밝힘.

핵심 기여도

핵심 아이디어

기존 연구는 크로스-토크나이저 온-포로시 디스틸레이션(OPD)에서 정렬 범위 확대가 학습 성능 향상에 기여한다고 가정했으나, 본 연구는 이 가정이 항상 성립하지 않음을 밝힘. 학생 모델이 생성한 토큰 중 85.57–96.98%는 1:1 정렬 그룹에 속해 있고, 이 위치에서 공통 토크나이저 단어가 대부분의 확률 질량을 차지함. 이는 정렬 범위가 넓다고 해서 반드시 더 많은 유용한 감독 정보를 얻는 것은 아니라는 통찰을 제공함. 학생이 선택한 상위 16개 단어로 reverse KL 손실을 제한하는 방식은 전체 공통 단어 기반 OPD와 유사한 정확도를 달성하면서도 기존 크로스-토크나이저 기반 기법보다 우수함. 이는 감독 신뢰도가 정렬 범위보다 더 중요한 요소임을 시사함.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 크로스-토크나이저 디스틸레이션에서 정렬 범위 확대보다 감독 신뢰도가 학습 성능에 더 큰 영향을 미친다는 점을 실증적으로 밝힘. 이는 디스틸레이션 연구에서 새로운 관점, 즉 감독 신뢰도 최적화를 위한 접근 방식을 제시함. 그러나 본 연구는 수학적 추론과 코드 생성에만 국한된 3개의 teacher–student 쌍을 대상으로 하였으며, 다른 도메인에서의 일반화 가능성은 추가 연구가 필요함. 또한, span 감독이 정확도를 낮추는 메커니즘에 대한 명확한 해석은 아직 부족함.

실용적 활용

본 연구의 결과는 토크나이저가 다른 모델 간의 지식 전이를 수행하는 데 활용 가능함. 특히, 정렬 범위 확대보다 감독 신뢰도를 높이는 방식이 더 효과적임을 보여주어, 디스틸레이션 모델의 성능 최적화에 기여할 수 있음. 코드 생성, 수학 문제 풀이 등 정밀도가 중요한 분야에서 활용 가능하며, 기존 크로스-토크나이저 기반 기법 대비 더 높은 성능을 보이는 top-k 기반 reverse KL 제한 방식을 실제 모델 학습에 적용할 수 있음.