dRAE: 超球面符号を用いた表現オートエンコーダ
dRAE: Representation Autoencoder with Hyper-Spherical Codes
July 24, 2026
著者: Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang, Junbo Zhao, Tong Zhang, Qixiang Ye
cs.AI
要旨
本研究では、高次元の視覚表現を離散化し、言語モデルとの橋渡しを図ることを目的とする。これは容易ではない課題であり、既存の量子化手法はコードブック崩壊を起こし、意味的一貫性を保ちながらスケーラビリティを実現できない。その根本原因は計量の不一致にある。すなわち、標準的なユークリッド空間上のコードブック目的関数は、表現空間の異方性幾何学と根本的に適合せず、その結果、コードブック埋め込みは分散の大きいマグニチュードスケールと不均一な角度分布を持ち、スケーラビリティを妨げる。この問題に対処するため、我々は超球面量子化(HSQ)を提案する。これは、角度に基づくルーティングによって意味内容と特徴量の大きさを分離し、コード割り当てが意味ではなくスケールに支配されるのを防ぐ。得られた離散表現オートエンコーダ(dRAE)は、高い忠実度で再構成を実現しつつ、意味的完全性を維持し、スケーラブルなコードブック予算をサポートする。広範な実験により、語彙サイズが131,072にスケールするにつれて一貫した性能向上が示され、加えて100%のコードブック利用率、簡素化された訓練パイプライン、理解および生成タスクにおける強力な性能が確認された。
English
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable codebook budget. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131{,}072, along with 100\% codebook utilization, simplified training pipeline, and strong performance across understanding and generation tasks.