Representation-Space MMD for Diffusion Language Models

Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev, Maksim Ignatov, Pavel Temirchev, Nikita Balagansky, Viacheslav Meshchaninov, Nikita Gushchin, Dmitry Baranchuk

arXiv:2610.06648 · 2026-10-06 공개 · arXiv · PDF

post-training gsm8k diffusion-language-models feature-space policy-gradients generative-perplexity mmd-estimation dmax-llada2-0

Abstract

We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for continuous models. In both cases, computing the loss directly from these features enables efficient post-training without full sampling trajectories or jointly trained auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, we increase decoding parallelism with similar or higher accuracy on math and code benchmarks.

한국어 요약

한 줄 요약

DLM의 생성 품질을 향상시키기 위해 MMD 최적화를 기반으로 post-training을 수행하는 방법을 제안한다.

핵심 기여도

핵심 아이디어

기존 DLM은 생성 과정에서 샘플링 경로나 별도의 보조 모델이 필요하지만, 본 연구는 생성된 토큰의 특징 공간에서 MMD를 직접 계산하여 post-training을 수행한다. 이는 기존의 복잡한 샘플링 과정이나 보조 모델 없이도 품질 향상을 가능하게 한다. MMD는 생성된 시퀀스와 참조 시퀀스의 특징 간 유사도를 측정하며, 이는 RBF 커널을 통해 계산된다. 이 연구는 DMax-LLaDA2.0 모델에서 hybrid masked-uniform diffusion 방식을 사용하여, 수학 및 코드 생성 벤치마크에서 정확도를 유지하면서 decoding parallelism을 증가시켰다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 DLM의 post-training을 효율적으로 수행할 수 있는 새로운 접근법을 제시하며, 별도의 보조 모델 없이도 생성 품질을 향상시킬 수 있음을 보여준다. 그러나 MMD 성능은 RBF bandwidth와 특징 공간 선택에 크게 의존하며, 이는 추가 연구가 필요한 부분이다. 또한, 단일 레이어의 특징만 사용하는 한계가 있으며, 다중 레이어나 노이즈 수준의 특징 결합이 성능 향상에 기여할 수 있다.

실용적 활용

본 방법은 대규모 DLM의 post-training을 효율적으로 수행할 수 있는 기반을 제공하며, 특히 수학적 추론 및 코드 생성과 같은 정확도가 중요한 분야에서 활용 가능하다. 또한, 별도의 보조 모델 없이도 품질 향상을 가능하게 하므로, 실시간 생성 시스템이나 자원 제한 환경에서도 유용할 수 있다.