News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow arXiv — NLP / Computation & Language research 3d ago PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning arXiv:2608.25486v1 Announce Type: cross Abstract: Long Narrative Reasoning is an essential capability for processing and reasoning over complex narratives. While retrieval-augmented generation provides a promising framework, existing methods still face two critical challenges:… 27 Hugging Face Daily Papers research 3d ago StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models Abstract StreamPI enhances single-frame vision-language-action models with streaming temporal reasoning via instruction-anchored attention and randomized interval training, improving robot manipulation without extra parameters. Generated by thinkingmachines/Inkling-Small… 6 Hugging Face Daily Papers research 3d ago V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning Abstract Visual Rubrics-Based Reinforcement Learning improves vision-language model grounding by scoring answers on visual faithfulness, reasoning consistency, and instruction following using structured partial credit. Generated by thinkingmachines/Inkling-Small Vision-language… 22 Hugging Face Daily Papers research 3d ago VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Abstract VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates. Generated by thinkingmachines/Inkling-Small Native visual reasoning treats visual generation as the… 6 Hugging Face Daily Papers research 3d ago VGI-BENCH: Probing Visual Intelligence in Video Generation Models Abstract VGI-bench evaluates visual reasoning in video generation models through 27 tasks, revealing limited reliability and minimal self-correction during generation. Generated by thinkingmachines/Inkling-Small Recent studies suggest that video generation models can exhibit… 25 Hugging Face Daily Papers research 3d ago Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments Abstract AnTrap benchmarks GUI agent robustness by injecting dynamic anomalies into execution trajectories, revealing universal vulnerabilities and distinguishing learnable traps from intrinsic reasoning limits. Generated by thinkingmachines/Inkling-Small GUI agents often… 19 r/LocalLLaMA community 3d ago N-gram vs Experts explained Since Qwen's dropped the Qwen4Exp architecture bomb that focus on offloading parameters to n-gram instead of pure mixture of experts, I dug into this and learned quite a lot. Here's the summary. Expect mistakes from human's writing lol. TLDR: MoEs do reasoning, N-grams do… 21 Vercel — AI dev-tools 3d ago Ling 3.0 Flash Fin now available on AI Gateway for free Ling 3.0 Flash Fin from Inclusion AI is now available on AI Gateway, free to use through September 25. Ling 3.0 Flash Fin is a finance-focused version of Ling 3.0 Flash . It has a 256K token context window, produces up to 32K output tokens, and supports reasoning and function… 29 r/MachineLearning community 3d ago A dataset with 52 Text to image model evaluation [P] I created a simple text to image benchmark. I curated 192 prompts that are difficult for T2I models in various ways: text rendering, spatial reasoning, human realism, negations, etc... I then asked a VLM to judge every output against a pre-specified binary question with the… 30 NVIDIA Developer Blog official-blog 3d ago NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,... 33 Hugging Face Daily Papers research 3d ago DREAM Technical Report Abstract DREAM introduces an agentic meta-control layer over industrial recommender pipelines that uses intent reasoning and dual-loop optimization to improve session-level recommendations without replacing existing models. Generated by thinkingmachines/Inkling-Small Industrial… 25 Hugging Face Daily Papers research 4d ago Meta^n: Recursive Self-Improvement through Emergent Depth Abstract Meta^n recursively applies a fixed meta-operation to growing inputs, building deeper reasoning layers that improve self-improving LLM agents without destabilizing the system. Generated by thinkingmachines/Inkling-Small Self-improving LLM agents refine answers, not the… 16 Hugging Face Daily Papers research 4d ago On-policy Distillation with Verifiable Reward Abstract OPDVR integrates on-policy distillation with verifiable rewards via a ReLU-gated implicit reward reformulation, improving reasoning performance without extra hyperparameters. Generated by thinkingmachines/Inkling-Small Reinforcement Learning with Verifiable Rewards… 35 Hugging Face Daily Papers research 4d ago Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Abstract OraRL improves reinforcement learning post-training for video multimodal language models by integrating oracle rollouts with decoupled advantage estimation and sign-balanced pruning, achieving higher sample efficiency and scalability without chain-of-thought generation.… 17 arXiv — Machine Learning research 4d ago Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders arXiv:2608.23809v1 Announce Type: new Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. We… 31 arXiv — Machine Learning research 4d ago Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents arXiv:2608.24087v1 Announce Type: new Abstract: Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own… 30 arXiv — Machine Learning research 4d ago Steering Recurrent Reasoners at Inference Time with Readout Feedback arXiv:2608.24136v1 Announce Type: new Abstract: Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more… 12 arXiv — Machine Learning research 4d ago When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study arXiv:2608.24492v1 Announce Type: new Abstract: Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings where ground-truth evidence is unavailable at inference time. Prior work has proposed… 17 r/LocalLLaMA community 4d ago ibm-granite/granite-4.2-30b · Hugging Face Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default),… 16 Hugging Face Daily Papers research 5d ago One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders Abstract Search-augmented LLM recommenders are highly vulnerable to web content polluted by generative engine optimization, frequently promoting fake products despite reasoning and defenses. Generated by thinkingmachines/Inkling-Small Search-augmented LLMs increasingly mediate… 26 Hugging Face Daily Papers research 5d ago Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion Abstract Block3D accelerates text-to-3D generation by using block-wise diffusion with confidence-guided correction to reduce inference time while preserving geometric fidelity. Generated by thinkingmachines/Inkling-Small While text-to-3D generation has advanced rapidly,… 13 arXiv — NLP / Computation & Language research 5d ago Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning arXiv:2608.21369v1 Announce Type: new Abstract: Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis,… 21 arXiv — NLP / Computation & Language research 5d ago MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning arXiv:2608.21808v1 Announce Type: new Abstract: Multimodal Retrieval-Augmented Generation (RAG) with visual citation is crucial for ensuring the traceability and verifiability of MLLMs. However, current RAG and SFT-based methods struggle to achieve robust cross-modal reasoning,… 18 arXiv — NLP / Computation & Language research 5d ago GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding arXiv:2608.21832v1 Announce Type: new Abstract: Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce… 37 arXiv — NLP / Computation & Language research 5d ago HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning arXiv:2608.21863v1 Announce Type: new Abstract: Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this… 16 arXiv — NLP / Computation & Language research 5d ago The Chase Is the Curriculum, the Capture Anchors the Credit: Pursuit-Evasion Self-Play for Zero-Data LLM Reasoning arXiv:2608.21871v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become the dominant recipe for improving large language model reasoning, yet it presumes large human-curated task collections. Zero-data self-play removes this dependency, but… 16 arXiv — NLP / Computation & Language research 5d ago Semantic Reasoning Denoising: Correcting Language Model Reasoning with Semantic Operators arXiv:2608.22090v1 Announce Type: new Abstract: Large language models can produce fluent reasoning traces whose local semantic errors propagate to an incorrect conclusion, while unconstrained self-correction may preserve, amplify, or introduce errors. Existing diffusion language… 20 arXiv — NLP / Computation & Language research 5d ago SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning arXiv:2608.22132v1 Announce Type: new Abstract: Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or… 7 arXiv — NLP / Computation & Language research 5d ago Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion arXiv:2608.22140v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong reasoning performance, but their robustness to realistic lexical corruption remains poorly understood. We evaluate four open-weight instruction-tuned models and frontier models across… 8 arXiv — NLP / Computation & Language research 5d ago Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching arXiv:2608.22332v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we… 15 arXiv — NLP / Computation & Language research 5d ago GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning arXiv:2608.22479v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge-intensive questions. For complex multi-hop questions, multi-turn retrieval-augmented reasoning extends RAG into an iterative… 30 arXiv — NLP / Computation & Language research 5d ago From Diagnosis to Redesign: Using Quantitative Ethnography to Improve Multi-Agent LLM Reasoning arXiv:2608.22566v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems are designed to improve reasoning by decomposing tasks across multiple agents with specialized functions, but the presence of multiple agents does not inherently guarantee coherent… 20 arXiv — NLP / Computation & Language research 5d ago Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains arXiv:2608.22622v1 Announce Type: new Abstract: Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language… 11 arXiv — NLP / Computation & Language research 5d ago Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a… 24 arXiv — NLP / Computation & Language research 5d ago SPOC-SQL: Stage-wise Preference Optimization for Controllable Text-to-SQL arXiv:2608.22772v1 Announce Type: new Abstract: Text-to-SQL aims to translate natural language questions into executable SQL queries over relational databases, requiring multi-stage structured reasoning over database schemas and query constraints. However, existing methods treat… 34 arXiv — NLP / Computation & Language research 5d ago DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation arXiv:2608.22806v1 Announce Type: new Abstract: Iterative preference optimization is essential for aligning Large Language Models on mathematical reasoning tasks, yet its efficiency is often throttled by signal scarcity: as the model improves, static problem sets become… 28 arXiv — NLP / Computation & Language research 5d ago SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning arXiv:2608.22857v1 Announce Type: new Abstract: Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names,… 20 Hugging Face Daily Papers research 5d ago Prime Agent: A Self-Improving RLM Harness Abstract Prime Agent is an open-source harness that uses recursive subagents, persistent computation, and agent-to-agent coordination to extend language models' long-horizon capabilities across coding and reasoning tasks. Generated by thinkingmachines/Inkling-Small Language… 22 Hugging Face Daily Papers research 5d ago Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress Abstract R2-OPD improves on-policy distillation by filtering teacher rewards that conflict with reasoning progress via within-trajectory ranking comparisons. Generated by thinkingmachines/Inkling-Small On-policy distillation (OPD) has emerged as an effective framework for… 25 r/LocalLLaMA community 5d ago Scaffold CoT: A CoT dataset built around the failures of small model (>5B Params) free form thinking. Hope its useful to you guys! TL;DR - A ~4M example, ~3B token CoT dataset designed around helping small models think more concisely, accurately and reliably. Hi all! For the past few months I have been working on a dataset designed around improving small model performance through a structured framework (or… 5 Hugging Face Daily Papers research 5d ago ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Abstract ParaTempo improves parallel reasoning efficiency by using temporal confidence to dynamically prune, retire, and reallocate reasoning branches without synchronization. Generated by thinkingmachines/Inkling-Small Parallel reasoning improves the accuracy and robustness of… 18 Hugging Face Daily Papers research 6d ago Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models Abstract On-policy distillation transfers reasoning behaviors rather than specific answers, with generalization strongly tied to teacher-student origin alignment and multi-teacher combinations causing capability trade-offs. Generated by thinkingmachines/Inkling-Small On-policy… 34 arXiv — NLP / Computation & Language research 6d ago Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck arXiv:2608.20362v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this… 33 arXiv — Machine Learning research 6d ago Harmonic Torsional Diffusion for Protein-Ligand Flexible Docking arXiv:2608.20366v1 Announce Type: cross Abstract: Molecular docking requires reasoning jointly about ligand pose and protein flexibility. Most diffusion-based docking models predict torsional updates with generic Euclidean heads that ignore the periodic geometry of angular… 18 arXiv — NLP / Computation & Language research 6d ago Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs arXiv:2608.20953v1 Announce Type: new Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding,… 5 arXiv — NLP / Computation & Language research 6d ago COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models arXiv:2608.21030v1 Announce Type: cross Abstract: Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal… 37 arXiv — NLP / Computation & Language research 6d ago When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these… 5 arXiv — NLP / Computation & Language research 6d ago Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing arXiv:2608.20348v1 Announce Type: new Abstract: Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than… 19 arXiv — NLP / Computation & Language research 6d ago Self-Speculation for Faster Reasoning Models arXiv:2608.20359v1 Announce Type: new Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor… 9 arXiv — NLP / Computation & Language research 6d ago ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research arXiv:2608.20391v1 Announce Type: new Abstract: Most legal NLP resources draw from federal case law and focus on coarse classification, leaving administrative adjudication, where the vast majority of government decisions occur, essentially unaddressed. We introduce… 10 Page 2 of 10 · 500 articles ← Newer Older →