News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow arXiv — NLP / Computation & Language research 12d ago $R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets arXiv:2608.16033v1 Announce Type: new Abstract: In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do… 12 arXiv — NLP / Computation & Language research 12d ago STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering arXiv:2608.16224v1 Announce Type: new Abstract: By leveraging large-scale pretraining, LLMs can interpret diverse temporal expressions and question formulations without task-specific training. However, existing prompt-based neuro-symbolic systems continue to rely on LLMs for… 11 arXiv — NLP / Computation & Language research 12d ago Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning arXiv:2608.16554v1 Announce Type: new Abstract: Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal:… 34 Hugging Face Daily Papers research 12d ago HarnessEval-W: Agentifying the Evaluation of Visual Worlds Abstract HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence. Generated by thinkingmachines/Inkling-Small A benchmark should deliver more than a scalar score: what makes an… 29 Hugging Face Daily Papers research 12d ago ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering Abstract ENTLORE is a benchmark framework that evaluates enterprise question answering by requiring recovery of implicit organizational relations across routine documents, revealing that even with gold sources many latent reasoning questions remain unanswered. Generated by… 7 Hugging Face Daily Papers research 12d ago R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets Abstract R³-Bench reveals that shared computation budgets cause reasoning agents to underperform relative to their single-problem capabilities across math, coding, and abstract reasoning tasks. Generated by thinkingmachines/Inkling-Small In cognitive science, resource… 7 Hugging Face Daily Papers research 12d ago NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents Abstract NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning. Generated by thinkingmachines/Inkling-Small… 17 r/LocalLLaMA community 12d ago Local agentic coding Benchmark : Qwen 3.8 27B (in many weights quants / cache quants / engine / reasoning effort) vs others. In medium reasoning mode, it both scores higher than the 3.6 version, AND is very much more efficient (almost half requests needed, and a third less tokens generated) - at DeepSeek v4 Flash 3107 MXFP4 level The xhigh mode is advertised to be the best one for hard tasks. In this… 32 r/LocalLLaMA community 12d ago Weirdly, no one talks about Temperature setting for the Qwen3.8 27b Mind you, it is 1.0 by default, yet everyone is focused on how much the new model thinks, restricting the reasoning budget and/or dropping the reasoning level. Set the temperature to 0.7 and the model will no longer write a whole book of thoughts before trying to make a small… 8 r/LocalLLaMA community 12d ago "Opus 4.8 thinks too much", "Muse Glimmer sits between Gemma and Qwen, that's boring", "Gemma 4 is too lazy" I'm starting to think there's no way to make a reasoning model that won't draw persistent vocal complaints on here. EDIT: Qwen 3.8 not Opus 4.8*, freudian slip lol   submitted by   /u/MerePotato [link]   [comments] 29 r/LocalLLaMA community 12d ago Qwen 3.8 27B Overthinking, It has to be done, it has to be overthinking to punch Opus 4.6 Yes, it sucks to waste time waiting on 16K+ reasoning tokens alone. But here's the thing, this is only a 27B model trying to perform on par with 1T+ parameter models. Something has to be sacrificed, and that sacrifice is the amount of reasoning or trajectory tokens. This isn't… 4 r/LocalLLaMA community 13d ago [Paper] Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory… 17 Hugging Face Daily Papers research 13d ago DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data Abstract Mimir v1 is a 1-billion-parameter Hierarchical Reasoning Model trained solely on permissible data that achieves competitive English results and state-of-the-art Danish performance across multiple benchmarks. Generated by thinkingmachines/Inkling-Small Current large… 20 r/LocalLLaMA community 13d ago Unpopular opinion : Qwen 3.8 27b is not an overthinker Yes it uses a ton more reasoning tokens than 3.6 did But test in on the same tasks with the other chinese models, glm 5.3, deepseek v4 flash and pro, etc it's really similar, and they are needed The reality is, we're just frustrated because our hardware do not allow most of us… 19 Hugging Face Daily Papers research 13d ago A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images Abstract The ALD/E-ImageMiner benchmark and ICDAR 2026 competition advance machine interpretation of scientific figures through tasks spanning visual reading, domain reasoning, and evidential justification, proposing long-term goals for verifiable multimodal scientific AI.… 30 Hugging Face Daily Papers research 13d ago Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning Abstract Mobius-v0 separates global memory storage from iterative reasoning modules to improve knowledge compression and inference efficiency, yielding comparable performance with less training data and faster inference. Generated by thinkingmachines/Inkling-Small We introduce… 24 Hugging Face Daily Papers research 13d ago SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation Abstract SPARGen unifies 3D reconstruction, dense correspondence, and spatial reasoning into a single instruction-conditioned multimodal generative model that jointly learns shared spatial representations. Generated by thinkingmachines/Inkling-Small Spatial perception and… 31 Hugging Face Daily Papers research 13d ago SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Abstract On-policy distillation from a long-context reasoning teacher to short-context students improves mathematical proof reasoning and generalizes to science benchmarks by aligning token spans, constraining length growth, and stabilizing training. Generated by… 26 Hugging Face Daily Papers research 13d ago Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Abstract Second Thought is a training-free framework that runs auxiliary reasoning branches in parallel during agent action-observation waits to reduce sequential decoding and turn counts without harming accuracy. Generated by thinkingmachines/Inkling-Small LLM agents in the… 35 Hugging Face Daily Papers research 13d ago MobileMem: Learning from a Year of Mobile Experiences Abstract MobileMem is a benchmark and framework for evaluating on-device long-term memory through year-scale, multimodal mobile experience trajectories that require temporal reasoning, knowledge updating, and preference inference. Generated by thinkingmachines/Inkling-Small The… 5 arXiv — NLP / Computation & Language research 13d ago Capacity-Dependent Effects of Data Selection for Reasoning arXiv:2608.13721v1 Announce Type: cross Abstract: In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that… 9 arXiv — Machine Learning research 13d ago More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It arXiv:2608.14420v1 Announce Type: new Abstract: Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time. It also has the potential to serve as a general-purpose front end… 6 arXiv — NLP / Computation & Language research 13d ago Modular Cognitive Architecture Emerges in Large Language Models arXiv:2608.13567v1 Announce Type: cross Abstract: The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular… 21 arXiv — NLP / Computation & Language research 13d ago Think in Latent, Explain in Language: Self-Explainable Latent Reasoning arXiv:2608.13570v1 Announce Type: new Abstract: Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing… 28 arXiv — NLP / Computation & Language research 13d ago GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings arXiv:2608.13698v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current… 23 arXiv — NLP / Computation & Language research 13d ago Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models arXiv:2608.13760v1 Announce Type: new Abstract: Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look… 22 arXiv — NLP / Computation & Language research 13d ago From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL arXiv:2608.13787v1 Announce Type: cross Abstract: AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. These principal-driven tasks routinely place the agent across from a counterpart (another… 15 arXiv — NLP / Computation & Language research 13d ago IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering arXiv:2608.13588v1 Announce Type: new Abstract: Multi-hop question answering requires complex reasoning across multiple evidence segments, which often overwhelms retrieval-augmented generation systems with lengthy and noisy contexts, thereby undermining both efficiency and… 18 arXiv — NLP / Computation & Language research 13d ago Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model arXiv:2608.14003v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference. In production settings, batched inference is… 11 arXiv — NLP / Computation & Language research 13d ago SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning arXiv:2608.14277v1 Announce Type: new Abstract: On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges,… 11 arXiv — NLP / Computation & Language research 13d ago You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model arXiv:2608.14465v1 Announce Type: new Abstract: A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper… 16 arXiv — NLP / Computation & Language research 13d ago Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation arXiv:2608.13712v1 Announce Type: cross Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator… 13 arXiv — NLP / Computation & Language research 13d ago Agentic Transaction: Towards ACID-Compliant Agent Systems arXiv:2608.13900v1 Announce Type: cross Abstract: Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation. As agents… 20 arXiv — NLP / Computation & Language research 13d ago Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages arXiv:2608.14375v1 Announce Type: cross Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a… 18 arXiv — NLP / Computation & Language research 13d ago Research-Oriented Human-Centric Evaluation for Foundation Models arXiv:2506.01793v2 Announce Type: replace Abstract: Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we… 14 arXiv — NLP / Computation & Language research 13d ago Adaptive Stopping for Multi-Turn LLM Reasoning arXiv:2604.01413v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve accuracy… 35 arXiv — NLP / Computation & Language research 13d ago Early Stopping for Large Reasoning Models via Confidence Dynamics arXiv:2604.04930v2 Announce Type: replace Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key challenge… 23 arXiv — NLP / Computation & Language research 13d ago BAT: Learning to Reason about Spatial Sounds with Large Language Models arXiv:2402.01591v4 Announce Type: replace-cross Abstract: Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural… 8 Hugging Face Daily Papers research 13d ago CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing Abstract CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance. Generated by thinkingmachines/Inkling-Small With the rapid advancement of… 36 Hugging Face Daily Papers research 13d ago Claim-Level Reliability Assessment for Efficient Test-Time Reasoning Abstract Claim-Level Reliability Assessment improves reasoning accuracy by verifying critical claims instead of sampling more solutions, reducing token use while boosting performance. Generated by thinkingmachines/Inkling-Small We propose claim-level falsification as a principle… 23 r/LocalLLaMA community 13d ago Anyone else get a kick out of Qwen 3.8 27B Reasoning Dialogue? I've been paying attention to the reasoning because I'm still evaluating the model, and I've just noticed that sometimes I get a kick out of the way this model's internal monologue seems to play out sometimes. Like I've seen it get genuinely frustrated with itself and express… 27 r/LocalLLaMA community 13d ago Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72 https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/ 4k tokens per second per GPU of which there are 72. 350 tokens per second per user "Without additional model tuning, the model achieves a… 13 r/LocalLLaMA community 14d ago Quick PSA: Qwen3.8-27B reasoning effort vs reasoning budget in llama.cpp If you are using llama-server with their web-ui for testing, keep in mind, that the reasoning selector is just a reasoning budget aka a hard cap and has, at least to my knowledge, nothing at all to do with Qwen3.8-27B's native reasoning effort capability! Selecting any value for… 34 r/LocalLLaMA community 14d ago Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute   submitted by   /u/juanviera23 [link]   [comments] 5 r/LocalLLaMA community 14d ago Qwen3.8 27B reasoning effort low/medium/xhigh comparison I did a short test of the different reasoning efforts, since on default xhigh the model thinks a lot . Not very scientific, just a quick "generate an SVG of a pelican on a bicycle" prompt with 3 different seeds. I think the result is interesting none the less: xhigh gives *much\… 28 r/LocalLLaMA community 14d ago Anyone managed to get Qwen 3.8 27B running smoothly on vLLM? Can't get rid of endless thinking Title pretty much says it all. I’ve deployed Qwen 3.8 27B using vLLM on an RTX 6000 Pro (tried multiple vLLM releases and launch recipes), but I can't get it into a usable state because of crazy long reasoning passes. Regardless of the thinking effort setting (xhigh, medium, or… 22 r/LocalLLaMA community 14d ago Gemma 4 E4B IQ2_XXS: + 140.54% Reasoning Performance From Tensor Level Quantization Allocation iq2_xxs tensor level allocation recovered reasoning from 28.9 -> 69.5 at the same 3.3gb budget. https://huggingface.co/ByteOtter/gemma-4-E4B-it-CADA-IQ2_XXS I posted my Gemma 4 12B q3 result a couple days ago, where tensor level allocation gave me an +8.55% relative improvement… 17 r/LocalLLaMA community 15d ago Try out this "high" reasoning mode for 27B (tested on VLLM) After a lot of tweaking, I have come to the conclusion that 27B lacks a reasoning mode that is between low and xhigh. The "medium" mode isn't actually medium, it erases the explicit instructions to the model. When medium is enabled, the model acts very differently - to me it… 20 r/MachineLearning community 15d ago BDH-CQ: IN-CONTEXT LEARNING WITH RECURRENT LATENT REASONING [R] We introduce BDH-CQ, a reasoning system that brings these capabilities together. Demonstrations of a previously unseen task update recurrent memory; the query is then solved through iterative computation in a high-dimensional latent workspace. Intermediate reasoning states are… 33 Hugging Face Daily Papers research 15d ago Thought-Level Beam Search for Reasoning Abstract Gambit improves reasoning model efficiency by using thought-level beam search to dynamically allocate compute to promising reasoning traces under fixed hardware budgets. Generated by thinkingmachines/Inkling-Small Test-time compute scaling is a primary driver of… 12 Page 5 of 10 · 500 articles ← Newer Older →