WeMM-Embedding is a family of universal multimodal embedding models developed by the WeChat Vision team. It provides unified representations for text, images, videos, visual documents, and interleaved multimodal inputs, achieving state-of-the-art performance across multiple benchmarks covering diverse tasks and domains.</p>\n","updatedAt":"2026-08-26T03:11:46.467Z","author":{"_id":"6564a2ceedae9c33b7654a1f","avatarUrl":"/avatars/42f09356a1282896573ccb44830cd327.svg","fullname":"JUNJIE ZHOU","name":"JUNJIE99","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":21,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png","fullname":"Tencent","name":"tencent","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8211896419525146},"editors":["JUNJIE99"],"editorAvatarUrls":["/avatars/42f09356a1282896573ccb44830cd327.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24053","authors":[{"_id":"6a8e57c47bc881afa25f30e6","name":"Junjie Zhou","hidden":false},{"_id":"6a8e57c47bc881afa25f30e7","name":"Ke Mei","hidden":false},{"_id":"6a8e57c47bc881afa25f30e8","name":"Lei Li","hidden":false},{"_id":"6a8e57c47bc881afa25f30e9","name":"Tianyi Wang","hidden":false},{"_id":"6a8e57c47bc881afa25f30ea","name":"Fengyun Rao","hidden":false},{"_id":"6a8e57c47bc881afa25f30eb","name":"Jing Lyu","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report","submittedOnDailyBy":{"_id":"6564a2ceedae9c33b7654a1f","avatarUrl":"/avatars/42f09356a1282896573ccb44830cd327.svg","isPro":false,"fullname":"JUNJIE ZHOU","user":"JUNJIE99","type":"user","name":"JUNJIE99"},"summary":"Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM-Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large-scale multimodal alignment stage, followed by a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Across extensive evaluations, WeMM-Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open-source baseline on MMEB-v2, while the 9B variant further achieves a new state-of-the-art overall score of 80.6. WeMM-Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e-commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM-Embedding.","upvotes":46,"discussionId":"6a8e57c57bc881afa25f30ec","githubRepo":"https://github.com/Tencent/WeMM-Embedding","githubRepoAddedBy":"user","ai_summary":"WeMM-Embedding is a family of universal multimodal embedding models that align text, images, videos, and interleaved inputs in a shared space, achieving state-of-the-art retrieval and recommendation performance across public benchmarks and large-scale WeChat applications.","ai_keywords":["multimodal embeddings","cross-scale knowledge transfer","fine-grained relevance supervision","MMEB-v2","universal multimodal embedding models"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":31,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6564a2ceedae9c33b7654a1f","avatarUrl":"/avatars/42f09356a1282896573ccb44830cd327.svg","isPro":false,"fullname":"JUNJIE ZHOU","user":"JUNJIE99","type":"user"},{"_id":"65ccbcfde75fc480714d1d04","avatarUrl":"/avatars/2def06230512892928537cf2d8499563.svg","isPro":false,"fullname":"Chaofan Li","user":"cfli","type":"user"},{"_id":"67d3eb5f96edf034dc522163","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67d3eb5f96edf034dc522163/Ije1Z1vXV3yZMylENd1iq.jpeg","isPro":false,"fullname":"JaysonCai","user":"Jayson236","type":"user"},{"_id":"6948e95a7c95c132e268530f","avatarUrl":"/avatars/c9d610d7db874c6d71b3cd79afba8dc4.svg","isPro":false,"fullname":"Benny Xie","user":"Tencentxie","type":"user"},{"_id":"667924007d43ca7ee5556d75","avatarUrl":"/avatars/d1a0f3cd3af0b882f2b496f7e8b233ab.svg","isPro":false,"fullname":"Qijie Wei","user":"QJWei","type":"user"},{"_id":"67067633351e0c16a5c27497","avatarUrl":"/avatars/356aa3431198c8931b820a714bcfb19d.svg","isPro":false,"fullname":"Shenghao Fu","user":"fushh7","type":"user"},{"_id":"664dabaade412305ff2c984e","avatarUrl":"/avatars/9dd55d8078f0534f2bdb8cfb84333b8e.svg","isPro":false,"fullname":"Zhang","user":"Daoze","type":"user"},{"_id":"6a4e03159872ca6c24dbbb09","avatarUrl":"/avatars/66eab3ebfb2b651d7c6c09dc680c80e6.svg","isPro":false,"fullname":"Bo Lin","user":"linb200","type":"user"},{"_id":"6145b3fd35135ec7e8d4ca45","avatarUrl":"/avatars/5dc25d18d6a8418c9b1a29ece9a48f5a.svg","isPro":false,"fullname":"Shuqi Lu","user":"shuqi","type":"user"},{"_id":"62d81d7dad693a1a9627b31d","avatarUrl":"/avatars/ff54cdce2b05ac1cd9128d033ef79748.svg","isPro":false,"fullname":"Guangting","user":"wgting96","type":"user"},{"_id":"63404ffc181f5648a58adcd3","avatarUrl":"/avatars/b7da03224ccf55c2e5eabc3786e8ae87.svg","isPro":false,"fullname":"K","user":"Rover00","type":"user"},{"_id":"65e54d95d821287383a220ed","avatarUrl":"/avatars/a22a6598ea7d50570194a4a7e5b5217d.svg","isPro":false,"fullname":"Lei","user":"ll490187880","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24053.md","query":{}}">
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
Abstract
WeMM-Embedding is a family of universal multimodal embedding models that align text, images, videos, and interleaved inputs in a shared space, achieving state-of-the-art retrieval and recommendation performance across public benchmarks and large-scale WeChat applications.
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems. In this report, we present WeMM-Embedding, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions. The family comprises 2B, 4B, and 9B variants and is trained in two stages: a large-scale multimodal alignment stage, followed by a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Across extensive evaluations, WeMM-Embedding achieves leading performance on multiple public benchmarks. Notably, the 2B variant already surpasses the previously leading 8B open-source baseline on MMEB-v2, while the 9B variant further achieves a new state-of-the-art overall score of 80.6. WeMM-Embedding also demonstrates strong practical performance across WeChat applications, with substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests. It has been deployed at scale across recommendation and search applications, including WeChat Channels, Official Accounts, Moments, and e-commerce services. We have released the model weights and code to facilitate future research at https://github.com/Tencent/WeMM-Embedding.
Community
WeMM-Embedding is a family of universal multimodal embedding models developed by the WeChat Vision team. It provides unified representations for text, images, videos, visual documents, and interleaved multimodal inputs, achieving state-of-the-art performance across multiple benchmarks covering diverse tasks and domains.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.24053 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.24053 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.