code: <a href=\"https://github.com/caipeng328/NaviDC-OCR\" rel=\"nofollow\">https://github.com/caipeng328/NaviDC-OCR</a><br>paper: <a href=\"https://arxiv.org/pdf/2608.12898\" rel=\"nofollow\">https://arxiv.org/pdf/2608.12898</a><br>huggingface: <a href=\"https://huggingface.co/StarDoc-AI/NaviDC-OCR\">https://huggingface.co/StarDoc-AI/NaviDC-OCR</a></p>\n","updatedAt":"2026-08-18T02:53:33.612Z","author":{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","fullname":"caipeng","name":"caipeng328","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8425590991973877},"editors":["caipeng328"],"editorAvatarUrls":["/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.12898","authors":[{"_id":"6a82d678b25f624fb96ac33a","user":{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","isPro":false,"fullname":"caipeng","user":"caipeng328","type":"user","name":"caipeng328"},"name":"Peng Cai","status":"claimed_verified","statusLastChangedAt":"2026-08-18T00:45:04.440Z","hidden":false},{"_id":"6a82d678b25f624fb96ac33b","name":"Zhaofan Zou","hidden":false},{"_id":"6a82d678b25f624fb96ac33c","name":"Shifa Liu","hidden":false},{"_id":"6a82d678b25f624fb96ac33d","name":"Yikun Wang","hidden":false},{"_id":"6a82d678b25f624fb96ac33e","name":"Jiawei Tang","hidden":false},{"_id":"6a82d678b25f624fb96ac33f","user":{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","isPro":false,"fullname":"Kaicheng Yang","user":"Kaichengalex","type":"user","name":"Kaichengalex"},"name":"Kaicheng Yang","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.084Z","hidden":false},{"_id":"6a82d678b25f624fb96ac340","name":"Meng Tong","hidden":false},{"_id":"6a82d678b25f624fb96ac341","name":"Zhongjiang He","hidden":false},{"_id":"6a82d678b25f624fb96ac342","name":"Hao Sun","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents","submittedOnDailyBy":{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","isPro":false,"fullname":"caipeng","user":"caipeng328","type":"user","name":"caipeng328"},"summary":"Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.","upvotes":4,"discussionId":"6a82d678b25f624fb96ac343","githubRepo":"https://github.com/caipeng328/NaviDC-OCR","githubRepoAddedBy":"user","ai_summary":"NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning.","ai_keywords":["Vision-Language Models","deformation-aware learning","adaptive sampling mechanism","content-structure decoupled learning","formula grammars","table structures","document parsing"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":17,"organization":{"_id":"6a7eafb9b293bf8f1ed9e4d9","name":"StarDoc-AI","fullname":"StarDoc-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/658976c5367c76b8ee314fc7/RRZ3jayAI1o2l5I8pkk5V.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"658976c5367c76b8ee314fc7","avatarUrl":"/avatars/529d99aed953cbbbabbe5d9e35eadae9.svg","isPro":false,"fullname":"caipeng","user":"caipeng328","type":"user"},{"_id":"63e202f352b7578dba448ab5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e202f352b7578dba448ab5/8itVBLcv14m7OVsoF8h1o.jpeg","isPro":false,"fullname":"Kaicheng Yang","user":"Kaichengalex","type":"user"},{"_id":"641d6cfa2797fc5011942524","avatarUrl":"/avatars/25eb223bc7739541b130c6fdd9399862.svg","isPro":false,"fullname":"liushifa","user":"liushifa","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a7eafb9b293bf8f1ed9e4d9","name":"StarDoc-AI","fullname":"StarDoc-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/658976c5367c76b8ee314fc7/RRZ3jayAI1o2l5I8pkk5V.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.12898.md","query":{}}">
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Abstract
NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning.
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.12898 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.12898 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.