WeMM-Embedding: 위챗 멀티모달 임베딩 기술 보고서
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
August 25, 2026
저자: Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu
cs.AI
초록
범용 멀티모달 임베딩은 현대 AI 시스템의 핵심 구성 요소로 자리 잡고 있으며, 이질적인 콘텐츠를 공유 공간에 표현하여 검색, 추천, 분류, 에이전트 시스템 등 다양한 애플리케이션을 지원한다. 본 보고서에서는 텍스트, 이미지, 비디오, 시각적 문서, 그리고 임의로 혼합된 멀티모달 입력을 지원하면서 유연한 출력 차원을 제공하는 범용 멀티모달 임베딩 모델 제품군인 WeMM-Embedding을 제시한다. 이 제품군은 2B, 4B, 9B 규모의 변형 모델로 구성되며, 대규모 멀티모달 정렬 단계와 그 후의 정제 단계라는 두 단계로 훈련된다. 정제 단계에서는 선별된 데이터, 세밀한 관련성 지도, 그리고 교차 규모 지식 전이를 활용한다. 광범위한 평가를 통해 WeMM-Embedding은 여러 공개 벤치마크에서 최고 수준의 성능을 달성한다. 특히 2B 변형 모델은 MMEB-v2에서 기존 최고 성능의 8B 오픈소스 기준 모델을 이미 능가하며, 9B 변형 모델은 종합 점수 80.6으로 새로운 최첨단 성능을 달성한다. WeMM-Embedding은 또한 WeChat 애플리케이션 전반에서 강력한 실용 성능을 입증하며, 26개 과제로 구성된 사내 벤치마크에서 상당한 성능 향상을 보이고 14개의 온라인 A/B 테스트에서 일관된 개선을 보여준다. 이 모델은 WeChat 채널, 공식 계정, 모멘트, 전자상거래 서비스를 포함한 추천 및 검색 애플리케이션에 대규모로 배포되었다. 향후 연구를 지원하기 위해 모델 가중치와 코드를 https://github.com/Tencent/WeMM-Embedding에 공개하였다.
English
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM-Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large-scale multimodal alignment stage, followed by a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Across extensive evaluations, WeMM-Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open-source baseline on MMEB-v2, while the 9B variant further achieves a new state-of-the-art overall score of 80.6. WeMM-Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e-commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM-Embedding.