WeMM-Embedding:微信多模态嵌入技术报告
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
August 25, 2026
作者: Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu
cs.AI
摘要
通用多模态嵌入正成为现代AI系统的核心组件,使异构内容能够在共享空间中进行表示,以支持检索、推荐、分类和智能体系统等应用。在本报告中,我们提出了WeMM-Embedding,一个通用多模态嵌入模型系列,支持文本、图像、视频、视觉文档以及任意交错的多模态输入,并具有灵活的输出维度。该系列包含2B、4B和9B三种变体,并采用两阶段训练:先是大规模多模态对齐阶段,然后是使用精选数据、细粒度相关性监督和跨尺度知识迁移的精调阶段。在广泛的评估中,WeMM-Embedding在多个公开基准上取得了领先性能。值得注意的是,2B变体在MMEB-v2上已超越此前领先的8B开源基线,而9B变体进一步取得了80.6的整体分数,达到新的最先进水平。WeMM-Embedding还在微信应用中展现出强大的实际性能,在26任务内部基准上取得了显著提升,并在14个在线A/B测试中持续改进。它已大规模部署于推荐和搜索应用,包括视频号、公众号、朋友圈和电商服务。我们已发布模型权重和代码,以促进未来研究,详见 https://github.com/Tencent/WeMM-Embedding。
English
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM-Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large-scale multimodal alignment stage, followed by a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Across extensive evaluations, WeMM-Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open-source baseline on MMEB-v2, while the 9B variant further achieves a new state-of-the-art overall score of 80.6. WeMM-Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e-commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM-Embedding.