KVAE:用於多模態生成模型的標記器系列
KVAE: Family of Tokenizers for Multimodal Generative Models
August 6, 2026
作者: Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov, Kirill Chernyshev, Kirill Malakhov, Ilia Vasiliev, Ilia Trushkin, Valeriya Kobenko, David Chikovani, Alexander Ivanov, Azat Saginbaev, Egor Silvestrov, Ivan Mikheev, Konstantin Zakharov
cs.AI
摘要
潛在擴散模型(LDM)是一種重要的範式,利用標記器將輸入訊號映射至壓縮表徵。這種依賴關係使標記器成為生成過程本身不可或缺的一部分,因為它會影響學習速度、合成樣本的品質,並為後續應用奠定基礎。本報告介紹了一系列適用於音訊、影像與影片的 KVAE 標記器,所有這些標記器皆專為後續的文字條件生成而設計:KVAE-Audio 是連續全頻帶 48 kHz 標記器,其潛在表示頻率為 50 Hz,包含 64 個通道;KVAE-3D 包含兩個因果影片標記器,分別支援 4x16x16 與 4x8x8 壓縮;KVAE-2D 是影像模型,以 32 個通道將輸入壓縮 8 倍。我們證明,重建結果(PSNR、LPIPS、PESQ 等)與生成結果在客觀指標(Frechet Distance、CLIP score、CLAP score 等)及主觀指標(並排評估)上,均達到或超越前沿的開源標記器,例如來自 Wan-2.2、HunyuanVideo-1.5、FLUX.2、MovieGen、StableAudio 與 MMAudio 的 VAE。考量到開發的難度,我們與社群分享訓練細節、模型選擇方法,以及對設計選擇的消融實驗。程式碼公開於 https://github.com/kandinskylab/kvae 與 https://github.com/kandinskylab/kvae-audio。
English
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-conditioned generation: KVAE-Audio, a continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE-3D -- two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE-2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side-by-side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at https://github.com/kandinskylab/kvae and https://github.com/kandinskylab/kvae-audio.