WeMM-Embedding:微信多模態嵌入技術報告
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
August 25, 2026
作者: Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu
cs.AI
摘要
通用多模態嵌入正逐漸成為現代AI系統的核心組成部分,使異質內容得以在共享空間中表示,從而支援檢索、推薦、分類及智能體系統等應用。在本報告中,我們提出WeMM-Embedding,一個通用多模態嵌入模型系列,支援文本、圖像、影片、視覺文件以及任意交錯的多模態輸入,並具備靈活的輸出維度。該系列包含2B、4B和9B三種規模的變體,並採用兩階段訓練:首先是大規模多模態對齊階段,其次是使用精選數據、細粒度相關性監督和跨尺度知識遷移的微調階段。在廣泛的評測中,WeMM-Embedding在多個公開基準上取得了領先的表現。值得注意的是,2B變體在MMEB-v2上已經超過了先前領先的8B開源基線,而9B變體進一步取得了80.6的整體新最佳成績。WeMM-Embedding在微信各類應用中也展現了強勁的實際效能,在包含26項任務的內部基準上取得顯著提升,並在14項線上A/B測試中表現出一致的改進。該模型已大規模部署於推薦和搜尋應用中,包括微信影片號、公眾號、朋友圈及電商服務。我們已開源模型權重和程式碼,以促進未來研究,網址為https://github.com/Tencent/WeMM-Embedding。
English
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM-Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large-scale multimodal alignment stage, followed by a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Across extensive evaluations, WeMM-Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open-source baseline on MMEB-v2, while the 9B variant further achieves a new state-of-the-art overall score of 80.6. WeMM-Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e-commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM-Embedding.