We introduce <strong>CLR (Claim-Level Reliability Assessment)</strong>, a training-free test-time scaling method previously used in <strong>VibeThinker-3B</strong>. CLR is built on a simple idea that exploits the asymmetry between solving and falsification to improve reasoning reliability.</p>\n<ol>\n<li><p><strong>Falsification exploits asymmetric capability requirements.</strong><br>With the same model parameters, establishing correctness through forward search is harder than falsifying a decisive claim. CLR exploits this asymmetry without requiring a stronger model.</p>\n</li>\n<li><p><strong>Claim-level verification improves the signal-to-noise ratio.</strong><br>By focusing on decision-critical claims, CLR reduces the influence of erroneous or irrelevant tokens in a full reasoning trace, making decisive failure signals easier to identify.</p>\n</li>\n<li><p><strong>Falsification compresses the survival space of incorrect reasoning.</strong><br>Rather than explicitly proving which trace is correct, CLR suppresses traces with decisive flaws, allowing more reliable reasoning to naturally emerge in the final consensus.</p>\n</li>\n</ol>\n<p>Empirically, this translates into substantial recovery of failed consensus. On GPT-OSS-20B, when at least one correct trace is already present but standard self-consistency still fails, CLR <strong>rescues ~37%</strong> of such cases on average. On CMIMC25, CLR outperforms standard self-consistency under matched model-call budgets, achieving <strong>82.19% vs. 77.50%</strong> while using <strong>37.0% fewer tokens,</strong> and delivers a <strong>+27.15</strong> pp gain over Pass@1.</p>\n<p>📄 Paper: <a href=\"https://arxiv.org/abs/2608.11994\" rel=\"nofollow\">https://arxiv.org/abs/2608.11994</a><br>💻 Code: <a href=\"https://github.com/WeiboAI/CLR\" rel=\"nofollow\">https://github.com/WeiboAI/CLR</a></p>\n","updatedAt":"2026-08-17T02:45:23.505Z","author":{"_id":"67486775ed2e4d9e50fc9117","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67486775ed2e4d9e50fc9117/WrCtPqY9X67ASbkUeloDF.jpeg","fullname":"Sen Xu","name":"SenXu1123","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8610867261886597},"editors":["SenXu1123"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/67486775ed2e4d9e50fc9117/WrCtPqY9X67ASbkUeloDF.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.11994","authors":[{"_id":"6a7d84da42823931a1f17359","user":{"_id":"67486775ed2e4d9e50fc9117","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67486775ed2e4d9e50fc9117/WrCtPqY9X67ASbkUeloDF.jpeg","isPro":false,"fullname":"Sen Xu","user":"SenXu1123","type":"user","name":"SenXu1123"},"name":"Sen Xu","status":"claimed_verified","statusLastChangedAt":"2026-08-13T16:45:05.093Z","hidden":false},{"_id":"6a7d84da42823931a1f1735a","name":"Wei Wang","hidden":false},{"_id":"6a7d84da42823931a1f1735b","name":"Shixi Liu","hidden":false},{"_id":"6a7d84da42823931a1f1735c","name":"Jixin Min","hidden":false},{"_id":"6a7d84da42823931a1f1735d","name":"Yingwei Dai","hidden":false},{"_id":"6a7d84da42823931a1f1735e","name":"Zhibin Yin","hidden":false},{"_id":"6a7d84da42823931a1f1735f","name":"Yirong Chen","hidden":false},{"_id":"6a7d84da42823931a1f17360","name":"Junlin Zhang","hidden":false}],"publishedAt":"2026-08-12T00:00:00.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"Claim-Level Reliability Assessment for Efficient Test-Time Reasoning","submittedOnDailyBy":{"_id":"67486775ed2e4d9e50fc9117","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67486775ed2e4d9e50fc9117/WrCtPqY9X67ASbkUeloDF.jpeg","isPro":false,"fullname":"Sen Xu","user":"SenXu1123","type":"user","name":"SenXu1123"},"summary":"We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verification. Since whole-trace evaluation often obscures decisive errors due to signal dilution from routine tokens, CLR condenses each reasoning trace into a compact set of decision-critical claims, thereby isolating its logical anchors. Furthermore, recognizing the inherent difficulty of generating entirely correct solutions under fixed model capabilities, CLR shifts the focus to semantic falsification. This approach exploits a fundamental asymmetry between solution construction and claim refutation. Constructing a valid solution requires a flawless reasoning path, whereas refuting an incorrect claim requires identifying only a single decisive flaw. This targeted search for negative evidence systematically compresses the survival space of high-confidence incorrect traces, effectively suppressing erroneous consensus via nonlinear reliability scoring. Across four LLMs and four reasoning benchmarks under matched budgets, CLR generally improves upon pass@1 and self-consistency. On GPT-OSS-20B/CMIMC25, for instance, CLR exceeds pass@1 by 27.15 percentage-points and raises self-consistency accuracy from 77.50\\% to 82.19\\% with 37.0\\% fewer tokens.","upvotes":9,"discussionId":"6a7d84db42823931a1f17361","githubRepo":"https://github.com/WeiboAI/CLR","githubRepoAddedBy":"user","ai_summary":"Claim-Level Reliability Assessment improves reasoning accuracy by verifying critical claims instead of sampling more solutions, reducing token use while boosting performance.","ai_keywords":["claim-level falsification","test-time scaling","Claim-Level Reliability Assessment","reasoning trace","semantic falsification","nonlinear reliability scoring","self-consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"68c8059479c43cfa50f36156","name":"WeiboAI","fullname":"WeiboAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64d1faaa1ed6649d70d1fa2f/lZVm6Yuiif9cdr5KsnfZr.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6a311ac85bf76b986db9256a","avatarUrl":"/avatars/47b37f684e79976eab5aaccf59cc088a.svg","isPro":false,"fullname":"zz1218","user":"zz1218","type":"user"},{"_id":"6a311b9d43f2af6bf52d98cf","avatarUrl":"/avatars/d8453166197b312eb40221c8e03670ae.svg","isPro":false,"fullname":"bjtuzz","user":"bjtuzz","type":"user"},{"_id":"6422eba32f38c0a50cfdc77d","avatarUrl":"/avatars/0b1137ff258ba578c8b0e257a43716fa.svg","isPro":false,"fullname":"lsx","user":"lsx666","type":"user"},{"_id":"6646fa14b096df5b522fd5f9","avatarUrl":"/avatars/98f67b6f8cb9776575916a2ba027f738.svg","isPro":false,"fullname":"MIN JIXIN","user":"JIXIN0121","type":"user"},{"_id":"68a6e9ac8258e261e1f63d75","avatarUrl":"/avatars/c299ba2255d40a9656dcbffdf9bae7b3.svg","isPro":false,"fullname":"xue","user":"linliaaa","type":"user"},{"_id":"64d1faaa1ed6649d70d1fa2f","avatarUrl":"/avatars/388ba18df077eaa8e16a89e59bf852fa.svg","isPro":false,"fullname":"YinZhiBin","user":"YinZhiBin","type":"user"},{"_id":"668b5090101353874ced73d0","avatarUrl":"/avatars/b2ec34a321890140e97ddd69884132a8.svg","isPro":false,"fullname":"junlin zhang","user":"junlinzhang","type":"user"},{"_id":"67486775ed2e4d9e50fc9117","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67486775ed2e4d9e50fc9117/WrCtPqY9X67ASbkUeloDF.jpeg","isPro":false,"fullname":"Sen Xu","user":"SenXu1123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68c8059479c43cfa50f36156","name":"WeiboAI","fullname":"WeiboAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64d1faaa1ed6649d70d1fa2f/lZVm6Yuiif9cdr5KsnfZr.png"},"query":{}}">
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
Published on Aug 12
· Submitted by Sen Xu on Aug 17 Abstract
Claim-Level Reliability Assessment improves reasoning accuracy by verifying critical claims instead of sampling more solutions, reducing token use while boosting performance.
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verification. Since whole-trace evaluation often obscures decisive errors due to signal dilution from routine tokens, CLR condenses each reasoning trace into a compact set of decision-critical claims, thereby isolating its logical anchors. Furthermore, recognizing the inherent difficulty of generating entirely correct solutions under fixed model capabilities, CLR shifts the focus to semantic falsification. This approach exploits a fundamental asymmetry between solution construction and claim refutation. Constructing a valid solution requires a flawless reasoning path, whereas refuting an incorrect claim requires identifying only a single decisive flaw. This targeted search for negative evidence systematically compresses the survival space of high-confidence incorrect traces, effectively suppressing erroneous consensus via nonlinear reliability scoring. Across four LLMs and four reasoning benchmarks under matched budgets, CLR generally improves upon pass@1 and self-consistency. On GPT-OSS-20B/CMIMC25, for instance, CLR exceeds pass@1 by 27.15 percentage-points and raises self-consistency accuracy from 77.50\% to 82.19\% with 37.0\% fewer tokens.
Community
We introduce CLR (Claim-Level Reliability Assessment), a training-free test-time scaling method previously used in VibeThinker-3B. CLR is built on a simple idea that exploits the asymmetry between solving and falsification to improve reasoning reliability.
Falsification exploits asymmetric capability requirements.
With the same model parameters, establishing correctness through forward search is harder than falsifying a decisive claim. CLR exploits this asymmetry without requiring a stronger model.
Claim-level verification improves the signal-to-noise ratio.
By focusing on decision-critical claims, CLR reduces the influence of erroneous or irrelevant tokens in a full reasoning trace, making decisive failure signals easier to identify.
Falsification compresses the survival space of incorrect reasoning.
Rather than explicitly proving which trace is correct, CLR suppresses traces with decisive flaws, allowing more reliable reasoning to naturally emerge in the final consensus.
Empirically, this translates into substantial recovery of failed consensus. On GPT-OSS-20B, when at least one correct trace is already present but standard self-consistency still fails, CLR rescues ~37% of such cases on average. On CMIMC25, CLR outperforms standard self-consistency under matched model-call budgets, achieving 82.19% vs. 77.50% while using 37.0% fewer tokens, and delivers a +27.15 pp gain over Pass@1.
📄 Paper: https://arxiv.org/abs/2608.11994
💻 Code: https://github.com/WeiboAI/CLR
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.11994 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.11994 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.11994 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.