r/LocalLLaMA · · 1 min read

tencent/WeMM-Embedding 9B/4B/2B

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported.

https://huggingface.co/tencent/WeMM-Embedding-9B

https://huggingface.co/tencent/WeMM-Embedding-4B

https://huggingface.co/tencent/WeMM-Embedding-2B

https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf

submitted by /u/jacek2023
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA