What happens when a language model must answer questions about a fixed document collection without receiving retrieved passages at inference time?</p>\n<p>We study this problem as <strong>document knowledge internalization</strong> and introduce <strong>IAR (Inject, Align, Recover)</strong>, a three-stage post-training framework:</p>\n<ul>\n<li><strong>Inject</strong> transforms documents into continuation, rewrite, and instruction-conditioned reconstruction objectives for dense knowledge exposure.</li>\n<li><strong>Align</strong> makes the injected knowledge accessible through answer-only QA supervision.</li>\n<li><strong>Recover</strong> merges the adapted checkpoint with the original instruction model to restore general capabilities.</li>\n</ul>\n<p>Across two document corpora and four model families—Llama, Phi, Qwen, and SmolLM—IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings. On average, it gains 3.6 percentage points in domain QA accuracy and 12.1 points in general performance across IFEval, MMLU, and MSBench.</p>\n<p>Our central takeaway is that document exposure, QA accessibility, and capability recovery should be measured and optimized separately rather than treated as a single fine-tuning problem. We welcome discussion on retrieval-free knowledge acquisition, evaluation protocols, and the trade-off between domain internalization and general capability retention.</p>\n","updatedAt":"2026-08-21T02:37:18.879Z","author":{"_id":"642f6c64f945a8a5c9ee5b5d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642f6c64f945a8a5c9ee5b5d/pHagRjAzTYHO4DCXCmBun.jpeg","fullname":"XiaofengShi","name":"XiaofengAlg","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8771790266036987},"editors":["XiaofengAlg"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642f6c64f945a8a5c9ee5b5d/pHagRjAzTYHO4DCXCmBun.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.20281","authors":[{"_id":"6a87affe89e517cbfd75dbe3","name":"Qian Kou","hidden":false},{"_id":"6a87affe89e517cbfd75dbe4","name":"Xiaofeng Shi","hidden":false},{"_id":"6a87affe89e517cbfd75dbe5","name":"Xiaosong Qiu","hidden":false},{"_id":"6a87affe89e517cbfd75dbe6","name":"Hua Zhou","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/642f6c64f945a8a5c9ee5b5d/Qx1_0iZ0RmlluIVHoyYZN.png","https://cdn-uploads.huggingface.co/production/uploads/642f6c64f945a8a5c9ee5b5d/YUtjPur1jw0CH68qW8JfS.png"],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization","submittedOnDailyBy":{"_id":"642f6c64f945a8a5c9ee5b5d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642f6c64f945a8a5c9ee5b5d/pHagRjAzTYHO4DCXCmBun.jpeg","isPro":false,"fullname":"XiaofengShi","user":"XiaofengAlg","type":"user","name":"XiaofengAlg"},"summary":"Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.","upvotes":3,"discussionId":"6a87affe89e517cbfd75dbe7","ai_summary":"IAR is a three-stage post-training framework that injects structured document knowledge into language models, aligns them for retrieval-free question answering, and recovers general capabilities, improving both domain accuracy and general performance.","ai_keywords":["document knowledge internalization","retrieval-free question answering","IAR","structured document knowledge injection","QA behavior alignment","general ability recovery","continuation rewrite instruction-conditioned reconstruction","LoRA","FAPM"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"61be9739d2f9358e24ca0a4f","name":"BAAI","fullname":"Beijing Academy of Artificial Intelligence","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1664511063789-632c234f42c386ebd2710434.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642f6c64f945a8a5c9ee5b5d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642f6c64f945a8a5c9ee5b5d/pHagRjAzTYHO4DCXCmBun.jpeg","isPro":false,"fullname":"XiaofengShi","user":"XiaofengAlg","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61be9739d2f9358e24ca0a4f","name":"BAAI","fullname":"Beijing Academy of Artificial Intelligence","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1664511063789-632c234f42c386ebd2710434.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.20281.md","query":{}}">
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
Abstract
IAR is a three-stage post-training framework that injects structured document knowledge into language models, aligns them for retrieval-free question answering, and recovers general capabilities, improving both domain accuracy and general performance.
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.
Community
What happens when a language model must answer questions about a fixed document collection without receiving retrieved passages at inference time?
We study this problem as document knowledge internalization and introduce IAR (Inject, Align, Recover), a three-stage post-training framework:
- Inject transforms documents into continuation, rewrite, and instruction-conditioned reconstruction objectives for dense knowledge exposure.
- Align makes the injected knowledge accessible through answer-only QA supervision.
- Recover merges the adapted checkpoint with the original instruction model to restore general capabilities.
Across two document corpora and four model families—Llama, Phi, Qwen, and SmolLM—IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings. On average, it gains 3.6 percentage points in domain QA accuracy and 12.1 points in general performance across IFEval, MMLU, and MSBench.
Our central takeaway is that document exposure, QA accessibility, and capability recovery should be measured and optimized separately rather than treated as a single fine-tuning problem. We welcome discussion on retrieval-free knowledge acquisition, evaluation protocols, and the trade-off between domain internalization and general capability retention.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.20281 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.20281 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.20281 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.