r/LocalLLaMA · · 1 min read

tencent/EVIE-Preview-4.5B · Hugging Face

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

tencent/EVIE-Preview-4.5B · Hugging Face

Overview

EVIE-Preview-4.5B is a state-of-the-art multilingual Visual Document Retrieval (VDR) model built upon Qwen3.5-4B. It employs ColBERT-style late interaction with native 128-dimensional multi-vector token embeddings (4.54B parameters, BF16).

By combining native GatedDeltaNet linear-attention and full-attention hybrid modeling with a compact visual projection, EVIE achieves top-tier performance across ViDoRe V1+V2 and ViDoRe V3 while generating ultra-compact 128D multi-vectors—cutting vector storage and indexing costs by 8× to 32× compared to wider 2560D–4096D representations.

Key Highlights

  • 🎯 Ultra-Compact 128D Index: Native 128-dimensional multi-vector representations drastically shrink downstream storage and index latency without sacrifice in retrieval precision.
  • 🏆 SOTA on ViDoRe Benchmarks: Outperforms larger 8B models on ViDoRe V3 (leading 7 of 8 public domains) and delivers top average accuracy on ViDoRe V1+V2 (85.93 nDCG@5).
  • 🌐 Robust Multilingual & Multi-Format: Strong zero-shot generalization across diverse languages (EN, FR, DE, IT, ES, PT, ZH, etc.) and visual formats (charts, tables, scientific reports, financial filings).
  • ⚡ Seamless ColPali Compatibility: Fully integrated with the standard colpali-engine ecosystem and late-interaction scoring pipelines.
submitted by /u/jacek2023
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA