KVAE: Family of Tokenizers for Multimodal Generative Models

Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov, Kirill Chernyshev, Kirill Malakhov, Ilia Vasiliev, Ilia Trushkin, Valeriya Kobenko, David Chikovani, Alexander Ivanov, Azat Saginbaev, Egor Silvestrov, Ivan Mikheev, Konstantin Zakharov

arXiv:2608.05798 · 2026-08-10 공개 · arXiv · PDF

latent-diffusion multimodal-model video-tokenizer clip-score psnr-metric vae-model text-conditioned-generation kvae-tokenizer

Abstract

Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-conditioned generation: KVAE-Audio, a continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE-3D -- two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE-2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side-by-side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at https://github.com/kandinskylab/kvae and https://github.com/kandinskylab/kvae-audio.

한국어 요약

한 줄 요약

KVAE는 텍스트 조건을 기반으로 생성하는 멀티모달 모델을 위한 새로운 토크나이저 시리즈로, 이미지, 비디오, 오디오에서 기존 오픈소스 토크나이저를 능가하는 성능을 보인다.

핵심 기여도

핵심 아이디어

KVAE는 텍스트 조건 생성 모델에서 레이턴트 공간의 품질과 압축률을 동시에 최적화하는 토크나이저를 제안한다. 기존 토크나이저는 레이턴트 공간의 'diffusability'와 압축률을 별도로 고려했지만, KVAE는 이를 통합적으로 설계하여 생성 속도와 품질을 동시에 향상시킨다. 특히, 비디오 토크나이저는 Conv3D를 주요 계산 블록으로 사용하며, 인과적 구조와 RMSNorm 정규화를 통해 효율적인 훈련과 추론을 가능하게 한다. 오디오 토크나이저는 48kHz의 고해상도 신호를 50Hz 레이턴트로 압축하면서도 PESQ, PSNR 등 재구성 지표에서 기존 모델을 능가한다.

기술적 접근법

주요 결과

의의 및 한계

KVAE는 멀티모달 생성 모델에서 레이턴트 토크나이저의 중요성을 강조하며, 다양한 압축 비율과 모델 구조를 통해 생성 품질과 속도를 균형 있게 향상시켰다. 특히, 디자인 선택에 대한 아블레이션 분석을 공개함으로써 토크나이저 개발의 가이드라인을 제시한다. 그러나, 높은 압축 비율은 레이턴트 공간의 정보 손실 가능성과 함께, 모델 복잡도 증가라는 한계가 존재한다.

실용적 활용

KVAE는 텍스트 기반 생성 모델을 활용하는 콘텐츠 제작 분야(예: 영상, 오디오, 이미지 생성)에서 즉각적으로 활용 가능하다. 특히, 고해상도 오디오 생성 및 비디오 편집 분야에서 기존 솔루션 대비 높은 성능과 효율성을 제공할 수 있다.