News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow Smol AI News news-outlet 10d ago not much happened today **OpenAI** and **Anthropic** expanded their agent platforms with new desktop features, collaborative editing, and composable APIs like Skills and Files API. **OpenAI** rolled out memory and workflow features in the EEA, UK, and Switzerland. **AT&T** revealed that 40% of employee… 17 arXiv — Machine Learning research 10d ago Towards Reversible Forgetting: Managing Obsolete Knowledge in Continual Enterprise AI Agents arXiv:2608.18177v1 Announce Type: new Abstract: Continual learning has traditionally treated forgetting as a failure, emphasizing preservation of previously acquired knowledge as environments evolve. We argue that this objective is incomplete for enterprise AI agents operating… 31 arXiv — Machine Learning research 10d ago Harness Continual Learning: Continual Adaptation Beyond Model Parameters arXiv:2608.19013v1 Announce Type: new Abstract: Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing… 11 arXiv — NLP / Computation & Language research 10d ago StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data arXiv:2608.18105v1 Announce Type: new Abstract: StocksTalk is a voice-enabled conversational system for transforming spoken financial screening requests into executable and validated structured queries over real-world market data. The system combines streaming speech… 9 arXiv — NLP / Computation & Language research 10d ago Persona-Guided LLM Agents for Task-Oriented Dialogue arXiv:2608.18085v1 Announce Type: new Abstract: Prior work has shown that large language models (LLMs) can express diverse personality traits in open-ended text generation. However, it remains unclear whether they can do so in a goal-directed dialogue without compromising task… 17 arXiv — NLP / Computation & Language research 10d ago DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models arXiv:2608.18103v1 Announce Type: new Abstract: Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are… 7 arXiv — NLP / Computation & Language research 10d ago Artifact-centered Claim-aware Observability for Autonomous Scientific Agents arXiv:2608.18312v1 Announce Type: new Abstract: Autonomous scientific agents now increasingly propose ideas, write code, run experiments, analyze results, and even draft papers. Observe and audit those agents are necessary but logging every model call is not enough, scientists… 24 arXiv — NLP / Computation & Language research 10d ago DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents arXiv:2608.18524v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, progress is fundamentally limited by the reliance on full-length trajectory imitation. For tasks… 30 arXiv — NLP / Computation & Language research 10d ago Beyond LLM-Based Reasoning: Lightweight GNNs for Agent Failure Attribution arXiv:2608.18575v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which frequently cause agents to produce incorrect outcomes. This motivates the task of Agent Failure Attribution: given a failed… 19 arXiv — NLP / Computation & Language research 10d ago MemFuse: Multi-Source Memory Fusion from Fragmented Observations arXiv:2608.18704v1 Announce Type: new Abstract: Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantly focus on single-source textual histories. In realistic settings, however, relevant… 10 arXiv — NLP / Computation & Language research 10d ago SPADE: Self-Play in Adaptive Synthetic Executable Environments arXiv:2608.19197v1 Announce Type: new Abstract: Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the… 10 arXiv — NLP / Computation & Language research 10d ago What Makes Software Issue Resolution Tasks Difficult for Agents? arXiv:2608.18280v1 Announce Type: cross Abstract: Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, benchmark scores remain difficult to interpret due to the lack of control and… 37 arXiv — NLP / Computation & Language research 10d ago ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents arXiv:2608.18307v1 Announce Type: cross Abstract: Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realistic component-centered interactions (e.g., toggle a… 26 arXiv — NLP / Computation & Language research 10d ago Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots arXiv:2608.18744v1 Announce Type: cross Abstract: Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying… 19 arXiv — NLP / Computation & Language research 10d ago What is Missing from AI Post-Training AI: An Empirical Analysis arXiv:2608.19072v1 Announce Type: cross Abstract: Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture… 13 Hugging Face Daily Papers research 10d ago SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation Abstract SemaPLC is a verification-gated agent harness that validates generated PLC logic through external compilation and live runtime execution, achieving higher verified pass rates than baseline methods. Generated by thinkingmachines/Inkling-Small Programmable logic… 20 Hugging Face Daily Papers research 10d ago FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Abstract FACET constructs executable terminal tasks by preserving source intent and grounding instructions, solutions, and verifiers in a shared repaired environment to enable scalable agent training. Generated by thinkingmachines/Inkling-Small Training terminal agents requires… 15 Vercel — AI dev-tools 10d ago Vercel Agent is now available in Slack code channels Vercel Agent now works in Slack code channels, a new kind of channel launched today for working with a coding agent. Anyone in the channel can follow the work, give Agent new instructions, and review the code it writes. Choose Create a code channel from the Slack sidebar, select… 38 NVIDIA Developer Blog official-blog 10d ago Developing NVIDIA Holoscan Applications with CLI, Skills, and AI Coding Agents NVIDIA Holoscan is a platform for building real-time AI applications at the edge, from medical imaging to robotics. HoloHub is its companion repository: a... 18 Anthropic SDK (Python) releases dev-tools 10d ago v0.125.0 0.125.0 (2026-08-19) Full Changelog: v0.124.0...v0.125.0 Features api: managed agents web search config and self hosted sandbox memory ( b75afd6 ) 37 Hacker News — AI on Front Page community 10d ago Feature Request: Support AGENTS.md Article URL: https://github.com/anthropics/claude-code/issues/6235 Comments URL: https://news.ycombinator.com/item?id=49367350 Points: 232 # Comments: 132 8 Hugging Face Daily Papers research 10d ago LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Abstract LEGO-RL connects native coding-agent harnesses to scalable policy-gradient training via in-process LLM proxying, sandbox orchestration, and integrated monitoring, improving sparse MoE model performance across multiple harnesses. Generated by… 33 NVIDIA Developer Blog official-blog 10d ago Evaluating AI Agent Skill Performance with NVIDIA SkillEvaluator AI agents are only as effective as the context they receive. Even with capable models and well-documented NVIDIA libraries, agents can spend extra steps finding... 18 r/LocalLLaMA community 11d ago Waiting for a 122B because of world knowledge? Any LLM will hallucinate the world knowledge, even a 3T model. Use a 4B with a kiwix skill and local Wikipedia, 50gb and no more hallucinated world knowledge. Ask your coding agent to build your own, with your rules and eventual fallback access to internet knowledge for what's… 29 r/LocalLLaMA community 11d ago Am I doing something wrong? Qwen 3.8 27B seems useless for agentic coding I have been using local models on/off for like 2 years or so but never really used them extensively because the closed ones were always much better. Once Qwen 3.8 27B was released I decided to give it another serious try. I configured Cline and ZooCode as VSCode addons,… 29 Hugging Face Daily Papers research 11d ago StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Abstract StartupBench evaluates end-to-end AI agents on real-world startup workflows and reveals that even top models complete only about 30% of tasks, highlighting gaps in instruction following and domain expertise. Generated by thinkingmachines/Inkling-Small Recent advances in… 7 Hugging Face Daily Papers research 11d ago aDSL: Agentic 3D Creation via Joint Agent-Program Design Abstract A co-designed domain-specific language and multi-agent system improve LLM-driven 3D program synthesis by using relational operators and iterative execution feedback. Generated by thinkingmachines/Inkling-Small Programmatic representations provide a compelling paradigm… 31 Hugging Face Daily Papers research 11d ago Demystifying Agent Skills: Why They Work-Until They Don't Abstract Skills enhance LLM agents primarily by stabilizing execution through procedural anchoring rather than injecting missing knowledge, though retrieval bottlenecks and brittle assumptions limit their effectiveness. Generated by thinkingmachines/Inkling-Small Skills have… 22 Hugging Face Daily Papers research 11d ago Agent Lightning v1.0: Towards Harnessed Agentic RL Abstract Agent Lightning v1.0 enables reproducible reinforcement learning for arbitrary agent harnesses, substantially improving coding-agent performance with minimal data and compute. Generated by thinkingmachines/Inkling-Small Modern agents operate inside agent harnesses that… 30 Hugging Face Daily Papers research 11d ago HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety Abstract HarnessRisk evaluates agent harness safety across six operational phases, revealing that configuration vulnerabilities and detection gaps allow high attack success despite preserved utility. Generated by thinkingmachines/Inkling-Small Large language models are… 31 arXiv — Machine Learning research 11d ago Agents unlock new capabilities through Switching LoRA Adapters as a Tool (SLAaaT) arXiv:2608.17034v1 Announce Type: new Abstract: Post-training can unlock new capabilities and improve performance on specialized tasks, but sometimes at the cost of catastrophic forgetting in other domains. This poses a problem in long agent trajectories that compose different… 6 arXiv — Machine Learning research 11d ago Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its… 25 arXiv — Machine Learning research 11d ago Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning arXiv:2608.17373v1 Announce Type: new Abstract: Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from high-dimensional visual inputs. Traditional sampling often relies on random or suboptimal… 4 arXiv — Machine Learning research 11d ago Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents arXiv:2608.17524v1 Announce Type: new Abstract: This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded… 31 arXiv — Machine Learning research 11d ago Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment arXiv:2608.17713v1 Announce Type: new Abstract: Agent evaluations and trace-based learning often compare outputs across transformed views through a post-response correspondence treated as neutral preprocessing. We show that this correspondence is a measurement intervention:… 33 arXiv — Machine Learning research 11d ago Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents arXiv:2608.18008v1 Announce Type: new Abstract: Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller… 33 arXiv — NLP / Computation & Language research 11d ago Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting arXiv:2608.17075v1 Announce Type: new Abstract: Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not… 28 arXiv — NLP / Computation & Language research 11d ago Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents arXiv:2608.17153v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can… 21 arXiv — NLP / Computation & Language research 11d ago Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback arXiv:2608.17587v1 Announce Type: new Abstract: Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution… 24 arXiv — NLP / Computation & Language research 11d ago CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion arXiv:2608.17911v1 Announce Type: new Abstract: As LLM agents operate across structured workflows and sessions, preserving long-term history does not ensure that later contexts can recover relevant evidence through a bounded memory interface. We study this evidence-reachability… 7 arXiv — NLP / Computation & Language research 11d ago SeqFeed: Improving Agentic RTL Code Generation with Sequential Behavior Feedback arXiv:2608.16934v1 Announce Type: cross Abstract: RTL code generation is a critical stage in hardware design, and the emergence of agentic systems offers new opportunities to automate this process. To generate correct RTL code, agents must understand sequential behavior,… 11 arXiv — NLP / Computation & Language research 11d ago Memory Is Communication: The Frontier Between Remembering and Signaling arXiv:2608.17053v1 Announce Type: cross Abstract: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks. Under… 27 arXiv — NLP / Computation & Language research 11d ago LLM-Derived Preference Judgments Are Not Self-Consistent arXiv:2608.17644v1 Announce Type: cross Abstract: Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work… 9 arXiv — NLP / Computation & Language research 11d ago On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification arXiv:2608.18066v1 Announce Type: cross Abstract: Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of… 8 arXiv — NLP / Computation & Language research 11d ago PrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference Training arXiv:2602.13840v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly deployed in personalized tasks involving sensitive, context-dependent information, where privacy violations may arise in agents' action due to the implicitness of contextual… 29 arXiv — NLP / Computation & Language research 11d ago DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling arXiv:2604.06474v2 Announce Type: replace Abstract: Deep research with Large Language Model (LLM) agents is emerging as a powerful paradigm for multi-step information discovery, synthesis, and analysis. However, existing approaches primarily focus on unstructured web data, while… 18 arXiv — NLP / Computation & Language research 11d ago SOD: Step-wise On-policy Distillation for Small Language Model Agents arXiv:2605.07725v3 Announce Type: replace Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model capacity. While reinforcement learning methods like group relative policy… 14 Hugging Face Daily Papers research 11d ago Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents Abstract Empirical evaluation of diverse memory substrates for long-horizon LLM agents reveals regime-dependent trade-offs, motivating adaptive substrate routing for reliable agent memory. Generated by thinkingmachines/Inkling-Small Memory is becoming core infrastructure for… 13 r/LocalLLaMA community 11d ago Ling-3.0-tiny is a very interesting model. Run on NVIDIA Orin Nano Super 8GB at 128K context with IQ4_NL quant. I have been searching for suitable model to run on my 8GB RAM toy, NVIDIA Orin Nano Super 8GB. This little toy was priced at $249 earlier this year (not any more), and pulls very little power when idle. It was an interesting device that suitable for an agent to host on. It is… 31 Hugging Face Daily Papers research 11d ago From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents Abstract RUPA models agent execution as a dependency graph to propagate uncertainty across long trajectories, improving failure detection and confidence estimation for LLM agents. Generated by thinkingmachines/Inkling-Small Reliable uncertainty quantification (UQ) is essential… 16 Page 6 of 10 · 500 articles ← Newer Older →