ChatPaper.aiChatPaper

WeMM-Embedding: WeChatマルチモーダルEmbedding技術報告

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

August 25, 2026
著者: Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu
cs.AI

要旨

ユニバーサルマルチモーダル埋め込みは、現代のAIシステムの中核的構成要素となりつつあり、異種コンテンツを共有空間で表現することで、検索、レコメンデーション、分類、エージェントシステムなどのアプリケーションを可能にしている。本レポートでは、テキスト、画像、動画、ビジュアル文書、任意にインターリーブされたマルチモーダル入力をサポートし、柔軟な出力次元を持つユニバーサルマルチモーダル埋め込みモデルのファミリーであるWeMM-Embeddingを紹介する。このファミリーは2B、4B、9Bのバリアントで構成され、大規模マルチモーダルアライメント段階と、厳選データ、細粒度の関連性に関する教師信号、クロススケール知識転移を用いたリファインメント段階の2段階で訓練される。広範な評価において、WeMM-Embeddingは複数の公開ベンチマークで最先端の性能を達成する。特に、2BバリアントはMMEB-v2において従来の最先端である8Bオープンソースベースラインをすでに上回り、9Bバリアントは総合スコア80.6という新たな最先端を達成する。WeMM-Embeddingはまた、WeChatアプリケーション群でも高い実用性能を示し、26タスクからなる社内ベンチマークで大幅な改善、14件のオンラインA/Bテストで一貫した改善をもたらした。本モデルは、WeChat Channels、公式アカウント、モーメンツ、ECサービスなど、レコメンデーションおよび検索アプリケーションに大規模に展開されている。将来の研究を促進するため、モデルの重みとコードを https://github.com/Tencent/WeMM-Embedding で公開している。
English
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM-Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large-scale multimodal alignment stage, followed by a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Across extensive evaluations, WeMM-Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open-source baseline on MMEB-v2, while the 9B variant further achieves a new state-of-the-art overall score of 80.6. WeMM-Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e-commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM-Embedding.