News / #reasoning Tag Reasoning 500 articles archived under #reasoning · RSS Sign in to follow r/LocalLLaMA community 14h ago This finance-model benchmark card is more useful for what it discloses than for who "wins" The official benchmark card for Ling-3.0-flash-Fin is a useful reminder that the unit being tested is rarely just “the model.” The release says most runs used temperature 1, top_p 0.95 and the highest available reasoning effort. FinFIRST and FinSearchComp Verified used a common… 22 Hugging Face Daily Papers research 1d ago CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes Abstract CritICL improves LLM reasoning at inference time by using structured failure patterns from weaker models as critique-based guidance, reducing generation and token costs. Generated by thinkingmachines/Inkling-Small Recent advances in inference-time scaling have… 9 r/MachineLearning community 1d ago New to this field need some guidance with my project( marine reasoning) [D] learnt about vector space , fields , and applications whatever i could then went on with learning python libraries like numpy , scikit learn , pandas , also had a bit of knowledge about tensorflow and how to use pytorch but now i stand so clueless when i try to apply my… 23 Hugging Face Daily Papers research 2d ago Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning Abstract An agentic framework combining LLMs and VLMs enables consistent, multi-instruction editing of long multi-shot videos while preserving spatiotemporal structure. Generated by thinkingmachines/Inkling-Small While generative AI has significantly advanced video editing,… 16 arXiv — Machine Learning research 2d ago GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion arXiv:2608.26585v1 Announce Type: new Abstract: Discrete diffusion models have become a strong, widely adopted class of generators for sequence data, and steering them toward a downstream reward at inference time, without any retraining, is increasingly important. Such… 24 arXiv — Machine Learning research 2d ago Performance Foundations of Parallel & Distributed Reasoning Language Models arXiv:2608.27046v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. The resulting recent Reasoning Language Models… 12 arXiv — Machine Learning research 2d ago Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO arXiv:2608.27351v1 Announce Type: new Abstract: Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to… 14 arXiv — NLP / Computation & Language research 2d ago AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking arXiv:2608.26141v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves performance on challenging tasks, current models… 13 arXiv — NLP / Computation & Language research 2d ago Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation,… 17 arXiv — NLP / Computation & Language research 2d ago TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault… 5 arXiv — NLP / Computation & Language research 2d ago Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound arXiv:2608.26136v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, and an appealing way to aim them at reasoning is to curate their data with a signal reinforcement learning already produces: the… 26 arXiv — NLP / Computation & Language research 2d ago CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models arXiv:2608.26147v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard… 21 arXiv — NLP / Computation & Language research 2d ago Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models arXiv:2608.26161v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently… 14 arXiv — NLP / Computation & Language research 2d ago Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification arXiv:2608.26329v1 Announce Type: new Abstract: While tool-augmented Large Language Models have significantly improved multi-step reasoning in quantitative STEM tasks, a critical residual failure mode remains: intermediate reasoning steps that are syntactically well-formed,… 17 arXiv — NLP / Computation & Language research 2d ago Co-Evolving Structured Knowledge and Reasoning in Language Models arXiv:2608.26386v1 Announce Type: new Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved… 23 arXiv — NLP / Computation & Language research 2d ago SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning arXiv:2608.26550v1 Announce Type: new Abstract: Reinforcement learning-based knowledge distillation has the potential to transfer complex reasoning from teacher to student models, yet it currently faces a critical dilemma: researchers must choose between sparse outcome-based… 34 arXiv — NLP / Computation & Language research 2d ago Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models arXiv:2608.26587v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) offer structured medical knowledge that can ground large language model (LLM) reasoning in clinical diagnosis application, yet how KG signal should be integrated into LLMs remains an open question.… 8 arXiv — NLP / Computation & Language research 2d ago Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD arXiv:2608.26735v1 Announce Type: new Abstract: Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities such as reasoning, coding, instruction following, and creative writing. We study this domain--general… 29 arXiv — NLP / Computation & Language research 2d ago RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models arXiv:2608.26832v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied to specialized domains, where effective use of domain expertise often requires reasoning over complex rules in concrete scenarios. However, existing benchmarks only partially… 5 arXiv — NLP / Computation & Language research 2d ago Reasoning about In-Context Samples for Machine-Translation arXiv:2608.27036v1 Announce Type: new Abstract: Large Language Models (LLMs) can be trained to perform chain-of-thoughts reasoning in order to improve the reliability of their responses. In this work, we investigate how explicit reasoning can be leveraged for LLM-Based Machine… 18 arXiv — NLP / Computation & Language research 2d ago When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue arXiv:2608.27176v1 Announce Type: new Abstract: Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or… 32 Hugging Face Daily Papers research 2d ago TTPO: Test-Time Policy Optimization Abstract Test-Time Policy Optimization enables label-free test-time training for mathematical reasoning by asymmetrically distilling agreeing rollouts and penalizing disagreeing ones, matching supervised performance. Generated by thinkingmachines/Inkling-Small Recent prominent… 6 Hugging Face Daily Papers research 2d ago Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning Abstract Aphanta evaluates when image-editing intermediates improve multimodal reasoning by testing direct, editor-generated, and idealized visual states across tasks. Generated by thinkingmachines/Inkling-Small Explicit visual intermediates can help multimodal large language… 26 Hugging Face Daily Papers research 2d ago CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension Abstract CaRGo-T improves multimodal humor understanding by modeling causal relationships as graph-based reasoning structures interpreted by vision-language models. Generated by thinkingmachines/Inkling-Small Large-scale vision-language models (VLMs) have demonstrated remarkable… 5 Hugging Face Daily Papers research 2d ago UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Abstract UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior. Generated by… 37 Hugging Face Daily Papers research 2d ago Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Abstract Evolution strategies improve reasoning diversity and Pass@K over GRPO through sparse functional updates and population diversity, supporting a hybrid training approach. Generated by thinkingmachines/Inkling-Small Evolution Strategies (ES) have recently emerged as a… 30 Vercel — AI dev-tools 2d ago Hy4 Preview now available on AI Gateway Hy4 Preview from Tencent is now available on AI Gateway. Hy4 Preview is an open-source Mixture-of-Experts model with 770B total parameters aimed at long-horizon coding, document analysis, game development, and scientific reasoning. It serves a context window of 1M tokens. To use… 32 r/LocalLLaMA community 2d ago Appreciation Post - thomsonreuters/Thomson-1.0-Small With the lack of support from Qwen regarding the smaller 9B and 35B MOE models. Like myself, not everyone is looking for an agentic coding model, I particularly use it for RAG and reviewing and require high reasoning across different documents & came across this Finetune:… 22 LangChain releases dev-tools 2d ago langchain-fireworks==1.6.1 Changes since langchain-fireworks==1.6.0 release(fireworks): 1.6.1 ( #39975 ) fix(fireworks): drop reasoning history blocks ( #39973 ) chore(model-profiles): refresh model profile data ( #39844 ) 4 Hugging Face Daily Papers research 2d ago A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans Abstract A modular medical imaging agent decomposes spatial relation verification into parsing, anatomical localization, and geometric rules to outperform end-to-end vision-language models on CT spatial reasoning. Generated by thinkingmachines/Inkling-Small Reliable spatial… 27 Hugging Face Daily Papers research 3d ago Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data Abstract Mixed supervised fine-tuning on combined reasoning corpora outperforms next-chunk reinforcement learning in efficiency and final accuracy across mathematical and out-of-domain tasks. Generated by thinkingmachines/Inkling-Small Recent work proposes next-chunk reasoning… 18 Hugging Face Daily Papers research 3d ago The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents Abstract Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once the hard reasoning… 37 arXiv — NLP / Computation & Language research 3d ago Demystifying Reinforcement Learning Post-Training of Language Models arXiv:2608.24949v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities. Yet for many researchers… 15 arXiv — Machine Learning research 3d ago AFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions arXiv:2608.24954v1 Announce Type: new Abstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication. We present AFDBench, an AI meteorologist that generates professional Area Forecast… 22 arXiv — NLP / Computation & Language research 3d ago Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference arXiv:2608.25542v1 Announce Type: cross Abstract: Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection… 18 arXiv — Machine Learning research 3d ago TailSFT: Filtered Fine-Tuning Improves Post-Training Performance arXiv:2608.25756v1 Announce Type: new Abstract: Reinforcement learning post-training drives reasoning and agentic capabilities in modern AI systems, yet a growing body of work shows that it is most effective when used to fine-tune an already capable base model. We question… 37 arXiv — NLP / Computation & Language research 3d ago Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation arXiv:2608.25277v1 Announce Type: new Abstract: Multi-agent LLM systems coordinate through natural-language messages that consume 40--60\% of their token budget. Replacing these with structured graphs reduces cost but fails on tasks requiring adaptive reasoning. We propose… 20 arXiv — NLP / Computation & Language research 3d ago Adaptive Triggering for Bias Correction in LLM Reasoning arXiv:2608.25379v1 Announce Type: new Abstract: Chain-of-thought prompting can expose and amplify demographic stereotypes within an LLM's intermediate reasoning and create a failure mode that final-answer debiasing alone cannot address. Mitigating such bias during generation… 11 arXiv — NLP / Computation & Language research 3d ago OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora arXiv:2608.25398v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a… 37 arXiv — NLP / Computation & Language research 3d ago DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models arXiv:2608.25428v1 Announce Type: new Abstract: Correcting flawed reasoning traces remains a significant challenge for Large Language Models (LLMs), whose autoregressive generation can propagate early mistakes into subsequent reasoning. We introduce DCGC, a Masked Diffusion… 15 arXiv — NLP / Computation & Language research 3d ago MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize arXiv:2608.25449v1 Announce Type: new Abstract: Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of… 5 arXiv — NLP / Computation & Language research 3d ago ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains arXiv:2608.25487v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information… 23 arXiv — NLP / Computation & Language research 3d ago ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives arXiv:2608.25531v1 Announce Type: new Abstract: Humanities and social science research requires close reading of long narrative materials such as novels, scripts, archives, and case reports, yet many users have limited access to costly proprietary long-context models. Compact,… 29 arXiv — NLP / Computation & Language research 3d ago GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning arXiv:2608.25583v1 Announce Type: new Abstract: Reasoning-oriented large language models often achieve strong problem-solving performance by generating long chains of thought, but this behavior substantially increases inference cost and latency. In contrast, instruction-tuned… 13 arXiv — NLP / Computation & Language research 3d ago AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification arXiv:2608.25637v1 Announce Type: new Abstract: Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable rewards. To improve verification accuracy, prior work has explored rule-based,… 20 arXiv — NLP / Computation & Language research 3d ago Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models arXiv:2608.25761v1 Announce Type: new Abstract: One common trade-off in the use of large language models involves reducing the size of the model while increasing the amount of computation at inference time, for example by using a wider beam search. In this paper, we examine the… 23 arXiv — NLP / Computation & Language research 3d ago Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation arXiv:2608.25922v1 Announce Type: new Abstract: GNN-based Knowledge Graph Question Answering (KGQA) pipelines process queries through four discrete stages: entity linking, subgraph retrieval, GNN reasoning, and answer generation. Standard robustness evaluations conflate… 15 arXiv — NLP / Computation & Language research 3d ago Prefix Sliding for efficient test-time scaling arXiv:2608.26070v1 Announce Type: new Abstract: Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that… 26 arXiv — NLP / Computation & Language research 3d ago PA-CoT: Profile-Adaptive Chain-of-Thought for Personalized Nutritional Consulting arXiv:2608.24907v1 Announce Type: cross Abstract: In health and nutrition consulting, widely used prompting methods pass the user profile as an unstructured block without a dedicated analysis step, leaving personalization as a critical structural gap. We introduce PA-CoT… 6 arXiv — NLP / Computation & Language research 3d ago Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace arXiv:2608.24958v1 Announce Type: cross Abstract: An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni… 8 Page 1 of 10 · 500 articles Older →