Hugging Face Daily Papers · · 6 min read

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual descriptions or localized visual regions as evidence proxies. However, such supervision signals may either overlook complex visual structures or provide incomplete and inaccurate representations of the underlying evidence. To address these limitations, we propose ConceptFormer, a latent concept representation learning framework for visual document retrieval. ConceptFormer models query-relevant evidence as continuous, query-conditioned latent concepts that explicitly bridge localized visual evidence and semantic relevance, without requiring either textual intermediate representations or direct reliance on raw visual annotations. During training, ConceptFormer employs a strong vision-language model to dynamically determine the number of latent concept tokens and uses these concepts as an intermediate representation to bridge the semantic gap between queries and documents, thereby guiding the learning of the embedding space. Experiments on diverse visual document retrieval benchmarks demonstrate that ConceptFormer achieves 16.7% and 22.1% relative improvements in average NDCG@10 over the strongest visual retrieval baseline and the strongest OCR-based text retrieval baseline, respectively. Further analysis reveals that latent concepts effectively connect localized visual evidence with semantic relevance, enabling the retriever to capture both fine-grained textual cues and complex document-level visual structures while preserving strong retrieval alignment.</p>\n","updatedAt":"2026-08-18T03:05:28.271Z","author":{"_id":"672b32daf801f80c9b697827","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/672b32daf801f80c9b697827/ywcKXA5EKBs9ryLiTQH75.png","fullname":"Chunyi Peng","name":"hmhm1229","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1670387859384-633fe7784b362488336bbfad.png","fullname":"OpenBMB","name":"openbmb","type":"org","isHf":false,"details":"Large Language Models","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8233332633972168},"editors":["hmhm1229"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/672b32daf801f80c9b697827/ywcKXA5EKBs9ryLiTQH75.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15698","authors":[{"_id":"6a83cb7c675db694db8cd4de","user":{"_id":"672b32daf801f80c9b697827","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/672b32daf801f80c9b697827/ywcKXA5EKBs9ryLiTQH75.png","isPro":false,"fullname":"Chunyi Peng","user":"hmhm1229","type":"user","name":"hmhm1229"},"name":"Peng Chunyi","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.149Z","hidden":false},{"_id":"6a83cb7c675db694db8cd4df","name":"Xu Zhipeng","hidden":false},{"_id":"6a83cb7c675db694db8cd4e0","name":"Yan Yukun","hidden":false},{"_id":"6a83cb7c675db694db8cd4e1","name":"Liu Zhenghao","hidden":false},{"_id":"6a83cb7c675db694db8cd4e2","name":"Yu Shi","hidden":false},{"_id":"6a83cb7c675db694db8cd4e3","name":"Mei Sen","hidden":false},{"_id":"6a83cb7c675db694db8cd4e4","name":"Sun Yubo","hidden":false},{"_id":"6a83cb7c675db694db8cd4e5","name":"Zhang Yongheng","hidden":false},{"_id":"6a83cb7c675db694db8cd4e6","name":"Zhou Jie","hidden":false},{"_id":"6a83cb7c675db694db8cd4e7","name":"Gu Yu","hidden":false},{"_id":"6a83cb7c675db694db8cd4e8","name":"Yu Ge","hidden":false},{"_id":"6a83cb7c675db694db8cd4e9","name":"Sun Maosong","hidden":false}],"publishedAt":"2026-08-16T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval","submittedOnDailyBy":{"_id":"672b32daf801f80c9b697827","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/672b32daf801f80c9b697827/ywcKXA5EKBs9ryLiTQH75.png","isPro":false,"fullname":"Chunyi Peng","user":"hmhm1229","type":"user","name":"hmhm1229"},"summary":"Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual descriptions or localized visual regions as evidence proxies. However, such supervision signals may either overlook complex visual structures or provide incomplete and inaccurate representations of the underlying evidence. To address these limitations, we propose ConceptFormer, a latent concept representation learning framework for visual document retrieval. ConceptFormer models query-relevant evidence as continuous, query-conditioned latent concepts that explicitly bridge localized visual evidence and semantic relevance, without requiring either textual intermediate representations or direct reliance on raw visual annotations. During training, ConceptFormer employs a strong vision-language model to dynamically determine the number of latent concept tokens and uses these concepts as an intermediate representation to bridge the semantic gap between queries and documents, thereby guiding the learning of the embedding space. Experiments on diverse visual document retrieval benchmarks demonstrate that ConceptFormer achieves 16.7\\% and 22.1\\% relative improvements in average NDCG@10 over the strongest visual retrieval baseline and the strongest OCR-based text retrieval baseline, respectively. Further analysis reveals that latent concepts effectively connect localized visual evidence with semantic relevance, enabling the retriever to capture both fine-grained textual cues and complex document-level visual structures while preserving strong retrieval alignment. Codes and data are available at https://github.com/Neuir/ConceptFormer.","upvotes":3,"discussionId":"6a83cb7d675db694db8cd4ea","githubRepo":"https://github.com/NEUIR/ConceptFormer","githubRepoAddedBy":"user","ai_summary":"ConceptFormer learns continuous latent concept representations to bridge visual evidence and semantic relevance for visual document retrieval without relying on text intermediates or raw visual annotations.","ai_keywords":["ConceptFormer","latent concept representation learning","vision-language model","latent concept tokens","embedding space","visual document retrieval","NDCG@10"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"672b32daf801f80c9b697827","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/672b32daf801f80c9b697827/ywcKXA5EKBs9ryLiTQH75.png","isPro":false,"fullname":"Chunyi Peng","user":"hmhm1229","type":"user"},{"_id":"65aa6ae215102fd65968615d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65aa6ae215102fd65968615d/Zs3ZXXblHZLEgVlawQV0p.jpeg","isPro":false,"fullname":"Yongheng Zhang","user":"BRZ911","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.15698.md","query":{}}">
Papers
arxiv:2608.15698

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

Published on Aug 16
· Submitted by
Chunyi Peng
on Aug 18
Authors:

Abstract

ConceptFormer learns continuous latent concept representations to bridge visual evidence and semantic relevance for visual document retrieval without relying on text intermediates or raw visual annotations.

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual descriptions or localized visual regions as evidence proxies. However, such supervision signals may either overlook complex visual structures or provide incomplete and inaccurate representations of the underlying evidence. To address these limitations, we propose ConceptFormer, a latent concept representation learning framework for visual document retrieval. ConceptFormer models query-relevant evidence as continuous, query-conditioned latent concepts that explicitly bridge localized visual evidence and semantic relevance, without requiring either textual intermediate representations or direct reliance on raw visual annotations. During training, ConceptFormer employs a strong vision-language model to dynamically determine the number of latent concept tokens and uses these concepts as an intermediate representation to bridge the semantic gap between queries and documents, thereby guiding the learning of the embedding space. Experiments on diverse visual document retrieval benchmarks demonstrate that ConceptFormer achieves 16.7\% and 22.1\% relative improvements in average NDCG@10 over the strongest visual retrieval baseline and the strongest OCR-based text retrieval baseline, respectively. Further analysis reveals that latent concepts effectively connect localized visual evidence with semantic relevance, enabling the retriever to capture both fine-grained textual cues and complex document-level visual structures while preserving strong retrieval alignment. Codes and data are available at https://github.com/Neuir/ConceptFormer.

Community

Paper author Paper submitter about 6 hours ago

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual descriptions or localized visual regions as evidence proxies. However, such supervision signals may either overlook complex visual structures or provide incomplete and inaccurate representations of the underlying evidence. To address these limitations, we propose ConceptFormer, a latent concept representation learning framework for visual document retrieval. ConceptFormer models query-relevant evidence as continuous, query-conditioned latent concepts that explicitly bridge localized visual evidence and semantic relevance, without requiring either textual intermediate representations or direct reliance on raw visual annotations. During training, ConceptFormer employs a strong vision-language model to dynamically determine the number of latent concept tokens and uses these concepts as an intermediate representation to bridge the semantic gap between queries and documents, thereby guiding the learning of the embedding space. Experiments on diverse visual document retrieval benchmarks demonstrate that ConceptFormer achieves 16.7% and 22.1% relative improvements in average NDCG@10 over the strongest visual retrieval baseline and the strongest OCR-based text retrieval baseline, respectively. Further analysis reveals that latent concepts effectively connect localized visual evidence with semantic relevance, enabling the retriever to capture both fine-grained textual cues and complex document-level visual structures while preserving strong retrieval alignment.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.15698
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.15698 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers