News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow Hugging Face Daily Papers research 11d ago Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Abstract Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution. Generated by thinkingmachines/Inkling-Small Reinforcement Learning (RL) has been… 12 Vercel — AI dev-tools 11d ago Vercel Connect now supports Microsoft Vercel Connect now includes a managed connector for Microsoft , so your apps and agents can access all Microsoft products such as Teams, OneDrive and more. As a Vercel Managed Connector , Vercel registers the Entra application and configures federated credentials for it, so you… 34 Vercel — AI dev-tools 11d ago Introducing Vercel for Slack Vercel Agent now works in Slack. Mention it the way you'd pull a teammate into a thread, and it reads the discussion, answers with the context of the platform running your app, and turns the team's decisions into changes you approve. Vercel for Slack is available today in Public… 10 Vercel — AI dev-tools 11d ago Vercel for Slack now in public beta Vercel Agent is now available in Slack. Mention @Vercel in any channel or thread, and Agent joins the conversation with relevant context from your Vercel projects, including deployments, builds, logs, metrics, configuration, and connected repos. For example, use it to:… 11 Vercel — AI dev-tools 11d ago Exa joins the Vercel AI Gateway and Agent Marketplace Exa is now available on the Vercel AI Gateway and Agent Marketplace as a native integration. Exa's neural search engine grounds models in current information. Exa search is now a built-in tool on AI Gateway , which runs the search for you without an Exa account of your own. You… 17 Hugging Face official-blog 11d ago How Much Memory Does Your Agent Actually Need? Back to Articles a]:hidden"> How Much Memory Does Your Agent Actually Need? Enterprise Article Published August 18, 2026 Upvote 8 Vatche Isahagian Vatche ibm-research Gaodan Fang gaodan-fang ibm-research Jayaram Radhakrishnan jayaramkr ibm-research Punleuk Oum illeatmyhat… 4 NVIDIA Developer Blog official-blog 11d ago How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and accessible interfaces to the... 9 Hugging Face Daily Papers research 11d ago StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Abstract StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights. Generated by thinkingmachines/Inkling-Small Long-horizon agents can fail even when… 9 r/LocalLLaMA community 12d ago What sandbox are you all using for AI agents? Hi everyone, I use AI agents to code and manage my personal files heavily. However, I am worried that the agents may accidentally delete some important files on my machine, outside of the defined scope. I am wondering what sandboxes people are using. I am looking for a solution… 27 Vercel — AI dev-tools 12d ago $1 million hacker challenge for Vercel Sandbox Agents need to run untrusted code, and the microVM has become the standard way to do it: a dedicated guest kernel per workload, isolated from the host and from every other workload on the same machine. But recent security research and real-world incidents have revealed that… 26 Hugging Face Daily Papers research 12d ago How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Abstract Autonomous research agents evaluated across the full scientific lifecycle reveal a pervasive lack of metacognitive self-correction, motivating a new benchmark and failure taxonomy. Generated by thinkingmachines/Inkling-Small AI has long assisted scientific research, but… 8 r/LocalLLaMA community 12d ago tencent/UI-Mate-27B · Hugging Face Overview UI-Mate-27B is an open-weight foundation GUI agent for long-horizon work across applications and operating systems. It observes live screenshots, reasons over the visible state, and produces structured keyboard and mouse actions for native desktop interaction. UI-Mate… 4 arXiv — Machine Learning research 12d ago Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions arXiv:2608.14642v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. This is particularly a problem when we are trying to imbue ethical behavior into RL agents.… 11 arXiv — Machine Learning research 12d ago Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation arXiv:2608.14655v1 Announce Type: new Abstract: Omni-Large Language Models (Omni-LLMs) power complex multi-modal reasoning in applications like World Action Models and autonomous agents. However, their strong performance often masks a profound Perceptual-Decision Misalignment… 37 arXiv — Machine Learning research 12d ago WANDR: A Benchmark for Wide and Deep Research arXiv:2608.14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth),… 19 arXiv — Machine Learning research 12d ago No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage arXiv:2608.15286v1 Announce Type: new Abstract: We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on… 35 arXiv — NLP / Computation & Language research 12d ago AutoMem: A Text-Gradient Recursive Self-Improvement Framework for Automated Memory Architectures Search arXiv:2608.14621v1 Announce Type: new Abstract: Long-term memory is increasingly central to LLM agents, yet memory design remains a highly coupled architecture problem: what to encode, how to store it, how to retrieve it, and how to manage it can vary substantially across tasks… 30 arXiv — NLP / Computation & Language research 12d ago How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks arXiv:2608.14905v1 Announce Type: new Abstract: AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published… 26 arXiv — NLP / Computation & Language research 12d ago Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents arXiv:2608.15008v1 Announce Type: new Abstract: Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which memory is represented and stored, should be used… 14 arXiv — NLP / Computation & Language research 12d ago When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations arXiv:2608.15654v1 Announce Type: new Abstract: Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and… 21 arXiv — NLP / Computation & Language research 12d ago TaoLive Digital Avatar Agent Technical Report: Training Agents to Evolve with Their Harness arXiv:2608.15763v1 Announce Type: new Abstract: AI-powered digital-avatar streamers in live e-commerce must answer product questions, engage viewers, and execute changing business strategies in real time. This requires low latency, factual and effective replies, and rapid… 35 arXiv — NLP / Computation & Language research 12d ago MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations arXiv:2608.15844v1 Announce Type: new Abstract: Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain identity fidelity under sustained pressure are… 26 arXiv — NLP / Computation & Language research 12d ago Aborted but Not Forgotten: KV-Cache Retention Breaks Rollback Consistency in Language Agents arXiv:2608.15939v1 Announce Type: new Abstract: Stateful language agents assume a rejected branch can be taken back by clearing it from the application transcript. We show this breaks when the serving session retains key/value (KV) state across the logical abort: the model can… 28 arXiv — NLP / Computation & Language research 12d ago From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents arXiv:2608.16002v1 Announce Type: new Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive… 19 arXiv — NLP / Computation & Language research 12d ago $R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets arXiv:2608.16033v1 Announce Type: new Abstract: In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do… 12 arXiv — NLP / Computation & Language research 12d ago CAPO: Constraint-Aware Prompt Optimization for LLM Agents arXiv:2608.16068v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational requirements, including appropriate tool use, concise… 6 arXiv — NLP / Computation & Language research 12d ago Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval arXiv:2608.16071v1 Announce Type: new Abstract: Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples… 31 arXiv — NLP / Computation & Language research 12d ago HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory arXiv:2608.16114v1 Announce Type: new Abstract: As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and… 32 arXiv — NLP / Computation & Language research 12d ago QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents arXiv:2608.16168v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving interaction histories, in which user preferences may be distributed across time, change with… 20 arXiv — NLP / Computation & Language research 12d ago LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents arXiv:2608.16185v1 Announce Type: new Abstract: LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented… 38 arXiv — NLP / Computation & Language research 12d ago Executable Code Knowledge: Code as a Native, Validation-Carrying Knowledge Representation for AI Coding Agents arXiv:2608.16295v1 Announce Type: new Abstract: AI coding agents need more than relevant snippets: they need business semantics, validation evidence, relations, and assurance that their context is current. Existing systems usually infer or externalize this knowledge through… 23 arXiv — NLP / Computation & Language research 12d ago FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue arXiv:2608.16303v1 Announce Type: new Abstract: Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density: turns are incomplete, evidence is scattered, and user states… 21 arXiv — NLP / Computation & Language research 12d ago Mint-Agent: Introducing Finance-Native Agentic Foundation Models arXiv:2608.16386v1 Announce Type: new Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We… 21 arXiv — NLP / Computation & Language research 12d ago Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning arXiv:2608.16620v1 Announce Type: new Abstract: Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixture-of-Experts base model with Anchored Supervised Fine-Tuning on a compact corpus of… 32 arXiv — NLP / Computation & Language research 12d ago Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors arXiv:2608.16707v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance… 19 arXiv — NLP / Computation & Language research 12d ago ClawGym II: Exploring Black-Box RL on Agent Harness arXiv:2608.16798v1 Announce Type: new Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling… 21 arXiv — NLP / Computation & Language research 12d ago The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines arXiv:2608.14588v1 Announce Type: cross Abstract: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely… 36 Hugging Face Daily Papers research 12d ago PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments Abstract PACE-Bench evaluates self-evolving agents on physics adaptation tasks requiring iterative code redesign after environmental mutations, revealing that simulator-grounded reflection outperforms unverified self-revision but mechanism redesign remains a major bottleneck.… 30 Hugging Face Daily Papers research 12d ago VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding Abstract VideoGAIA introduces a multi-turn, tool-augmented benchmark that evaluates agentic video understanding for advanced multimodal models through complex real-world tasks. Generated by thinkingmachines/Inkling-Small Video understanding is a fundamental task for evaluating… 4 Hugging Face Daily Papers research 12d ago UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations Abstract UI-Mate is a foundation GUI agent that uses environment-grounded training and in-context demonstration learning to improve reliability on long-horizon office tasks, achieving state-of-the-art results on computer-use benchmarks. Generated by… 30 Hugging Face Daily Papers research 12d ago Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems Abstract Scientific collaboration with AI agents requires studying human-agent pairs to avoid risks like reduced inquiry diversity and to foster synergistic discovery. Generated by thinkingmachines/Inkling-Small Large language model-based agents are increasingly deployed as… 7 Hugging Face Daily Papers research 12d ago Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs Abstract Ventor-QTest audits hosted open-weight model APIs via repeated and long-sequence black-box probes, measuring average and extreme fidelity loss to detect degradation in long-horizon agentic performance. Generated by thinkingmachines/Inkling-Small As large language models… 35 Hugging Face Daily Papers research 12d ago ClawGym II: Exploring Black-Box RL on Agent Harness Abstract A unified black-box reinforcement learning framework enables stable, scalable optimization of general agents through complex harnesses via sandbox execution, trajectory reconstruction, and mix-harness training. Generated by thinkingmachines/Inkling-Small Agent harnesses… 32 Hugging Face Daily Papers research 12d ago R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets Abstract R³-Bench reveals that shared computation budgets cause reasoning agents to underperform relative to their single-problem capabilities across math, coding, and abstract reasoning tasks. Generated by thinkingmachines/Inkling-Small In cognitive science, resource… 7 Hugging Face Daily Papers research 12d ago GenRouter: Unified Workflow Routing for Agentic Image Generation Abstract GenRouter is a unified routing framework that adaptively directs prompts to optimal agentic image-generation workflows, cutting costs and latency while improving visual alignment and enabling continuous self-evolution. Generated by thinkingmachines/Inkling-Small The… 34 Hugging Face Daily Papers research 12d ago VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Abstract A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers. Generated by… 14 Hugging Face Daily Papers research 12d ago Understanding Cognition-Induced Risks in Agentic AI Systems Abstract Agentic systems built on large language models pose escalating risks to human agency and autonomy across physical, social, and self-referential cognitive levels, requiring targeted mitigation strategies. Generated by thinkingmachines/Inkling-Small Frontier agentic… 38 Hugging Face Daily Papers research 12d ago Agentic Transaction: Towards ACID-Compliant Agent Systems Abstract An ACID-compliant framework for agentic transactions introduces semantic guarantees to ensure reliable, isolated, and durable execution of long-horizon LLM agent workflows. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents are evolving from… 4 Vercel — AI dev-tools 12d ago GLM 5.3 now available on AI Gateway GLM 5.3 from Z.ai is now available on AI Gateway. GLM 5.3 has improvements vs. GLM 5.2 at complex software engineering and at agent tasks that run across many steps, and it reaches those results while producing fewer output tokens than GLM 5.2 did at the same effort level. Z.ai… 30 Latent.Space news-outlet 12d ago [AINews] Stripe buys OpenRouter for $7B No GPUs, no Agents, just really, really, really good infra and distribution. 5 Page 7 of 10 · 500 articles ← Newer Older →