Hugging Face Daily Papers · · 3 min read

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

code: <a href=\"https://github.com/caipeng328/NaviDC-OCR\" rel=\"nofollow\">https://github.com/caipeng328/NaviDC-OCR</a><br>paper: <a href=\"https://arxiv.org/pdf/2608.12898\" rel=\"nofollow\">https://arxiv.org/pdf/2608.12898</a><br>huggingface: <a href=\"https://huggingface.co/StarDoc-AI/NaviDC-OCR\">https://huggingface.co/StarDoc-AI/NaviDC-OCR</a></p>\n","updatedAt":"2026-08-18T02:53:33.612Z","author":{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","fullname":"caipeng","name":"caipeng328","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8425590991973877},"editors":["caipeng328"],"editorAvatarUrls":["/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.12898","authors":[{"_id":"6a82d678b25f624fb96ac33a","user":{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","isPro":false,"fullname":"caipeng","user":"caipeng328","type":"user","name":"caipeng328"},"name":"Peng Cai","status":"claimed_verified","statusLastChangedAt":"2026-08-18T00:45:04.440Z","hidden":false},{"_id":"6a82d678b25f624fb96ac33b","name":"Zhaofan Zou","hidden":false},{"_id":"6a82d678b25f624fb96ac33c","name":"Shifa Liu","hidden":false},{"_id":"6a82d678b25f624fb96ac33d","name":"Yikun Wang","hidden":false},{"_id":"6a82d678b25f624fb96ac33e","name":"Jiawei Tang","hidden":false},{"_id":"6a82d678b25f624fb96ac33f","user":{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","isPro":false,"fullname":"Kaicheng Yang","user":"Kaichengalex","type":"user","name":"Kaichengalex"},"name":"Kaicheng Yang","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.084Z","hidden":false},{"_id":"6a82d678b25f624fb96ac340","name":"Meng Tong","hidden":false},{"_id":"6a82d678b25f624fb96ac341","name":"Zhongjiang He","hidden":false},{"_id":"6a82d678b25f624fb96ac342","name":"Hao Sun","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents","submittedOnDailyBy":{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","isPro":false,"fullname":"caipeng","user":"caipeng328","type":"user","name":"caipeng328"},"summary":"Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.","upvotes":4,"discussionId":"6a82d678b25f624fb96ac343","githubRepo":"https://github.com/caipeng328/NaviDC-OCR","githubRepoAddedBy":"user","ai_summary":"NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning.","ai_keywords":["Vision-Language Models","deformation-aware learning","adaptive sampling mechanism","content-structure decoupled learning","formula grammars","table structures","document parsing"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":17,"organization":{"_id":"6a7eafb9b293bf8f1ed9e4d9","name":"StarDoc-AI","fullname":"StarDoc-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/658976c5367c76b8ee314fc7/RRZ3jayAI1o2l5I8pkk5V.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","isPro":false,"fullname":"caipeng","user":"caipeng328","type":"user"},{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","isPro":false,"fullname":"Kaicheng Yang","user":"Kaichengalex","type":"user"},{"_id":"641d6cfa2797fc5011942524","avatarUrl":"/avatars/25eb223bc7739541b130c6fdd9399862.svg","isPro":false,"fullname":"liushifa","user":"liushifa","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a7eafb9b293bf8f1ed9e4d9","name":"StarDoc-AI","fullname":"StarDoc-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/658976c5367c76b8ee314fc7/RRZ3jayAI1o2l5I8pkk5V.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.12898.md","query":{}}">
Papers
arxiv:2608.12898

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Published on Aug 13
· Submitted by
caipeng
on Aug 18
Authors:

Abstract

NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning.

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.12898
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.12898 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.12898 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers