dRAE: 具有超球面编码的表示自编码器
dRAE: Representation Autoencoder with Hyper-Spherical Codes
July 24, 2026
作者: Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang, Junbo Zhao, Tong Zhang, Qixiang Ye
cs.AI
摘要
本文旨在对高维视觉表示进行离散化,以弥合其与语言模型之间的鸿沟——这一挑战颇具难度,因为现有量化方法存在码本坍缩问题,在扩展时难以保持语义连贯性。我们发现根本原因在于度量失配:标准的欧几里得码本目标函数与表示空间的各向异性几何结构本质不匹配,导致码本嵌入产生高方差幅度尺度和不均匀的角度分布,从而阻碍可扩展性。为解决此问题,我们提出超球面量化(HSQ),该方法通过角度路由将语义内容与特征幅度解耦,避免码分配被尺度而非语义主导。由此得到的离散表示自编码器(dRAE)在保持语义完整性和支持可扩展码本预算的同时,实现了高保真重建。大量实验表明,当词汇表规模扩展至131,072时,该方法持续提升性能,同时实现100%码本利用率、简化训练流程,并在理解与生成任务中均展现出强劲表现。
English
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable codebook budget. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131{,}072, along with 100\% codebook utilization, simplified training pipeline, and strong performance across understanding and generation tasks.