dRAE:具有超球面編碼的表示自編碼器
dRAE: Representation Autoencoder with Hyper-Spherical Codes
July 24, 2026
作者: Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang, Junbo Zhao, Tong Zhang, Qixiang Ye
cs.AI
摘要
在本研究中,我們旨在對高維視覺表徵進行離散化,以彌合其與語言模型之間的差距——這是一項非平凡的挑戰,因為現有量化方法面臨碼簿崩潰問題,無法在保持語義一致性的同時進行擴展。我們發現根本原因在於度量不匹配:標準的歐幾里得碼簿目標函數從根本上與表徵空間的各向異性幾何不一致,導致碼簿嵌入產生高方差量值尺度與不均勻的角度分佈,進而阻礙可擴展性。為解決此問題,我們提出超球面量化(HSQ),透過角度路由將語義內容與特徵量值解耦,避免碼分配受到尺度而非意義的支配。所產生的離散表徵自編碼器(dRAE)在保持語義完整性並支援可擴展碼簿預算的同時,實現了高保真重建。大量實驗顯示,隨著詞彙規模擴展至131,072,效能持續提升,同時達成100%碼簿利用率、簡化的訓練流程,並在理解與生成任務中均展現優異表現。
English
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable codebook budget. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131{,}072, along with 100\% codebook utilization, simplified training pipeline, and strong performance across understanding and generation tasks.