Hugging Face Daily Papers · · 3 min read

Automata from Agent Traces: Failure and Next-Step Prediction

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

You write Agent State Machine and it drives the program.<br>This is the inverse: take a corpus of agent traces, recover the machine that produced them.<br>The machine belongs to the harness, not the LLM.</p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/64105805928400b416439f10/CAaswF8FGlReDAooYT-t7.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-08-26T23:19:24.247Z","author":{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","fullname":"Seonglae Cho","name":"seonglae","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.9164570569992065},"editors":["seonglae"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23670","authors":[{"_id":"6a8efd672c24e8c5fab326ac","user":{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae","type":"user","name":"seonglae"},"name":"Seonglae Cho","status":"claimed_verified","statusLastChangedAt":"2026-08-26T16:45:04.654Z","hidden":false},{"_id":"6a8efd672c24e8c5fab326ad","name":"Franklin Cardenoso Fernandez","hidden":false},{"_id":"6a8efd672c24e8c5fab326ae","name":"Umar Mohammed","hidden":false},{"_id":"6a8efd672c24e8c5fab326af","name":"Zekun Wu","hidden":false},{"_id":"6a8efd672c24e8c5fab326b0","name":"Kleyton Da Costa","hidden":false},{"_id":"6a8efd672c24e8c5fab326b1","name":"Ilham Wicaksono","hidden":false},{"_id":"6a8efd672c24e8c5fab326b2","name":"Adriano Koshiyama","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"Automata from Agent Traces: Failure and Next-Step Prediction","submittedOnDailyBy":{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae","type":"user","name":"seonglae"},"summary":"LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.","upvotes":4,"discussionId":"6a8efd672c24e8c5fab326b3","projectPage":"https://seongland.com/article/asg","ai_summary":"LLM agent traces are compressed into compact finite-state machines that enable accurate next-step and failure prediction for safety auditing and runtime monitoring.","ai_keywords":["finite-state machine","LLM agents","behavioral topology","next-step prediction","failure prediction","runtime monitoring","safety auditing","cross-run topology"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"64a32e214c758629715d2607","name":"holistic-ai","fullname":"Holistic AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648c6ae855101b9f9affdbc7/MFj_bUiWw-TFpouPKQian.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae","type":"user"},{"_id":"68383cfda2d6f83cf1ee88a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/mGNmuxzz1zaPZyMQFCde6.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae-holistic","type":"user"},{"_id":"631e14ac473a6825f285e89d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/631e14ac473a6825f285e89d/K-6QnoeGLg8XFvbTMMdqA.jpeg","isPro":false,"fullname":"Yury Panikov","user":"panikov","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64a32e214c758629715d2607","name":"holistic-ai","fullname":"Holistic AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648c6ae855101b9f9affdbc7/MFj_bUiWw-TFpouPKQian.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23670.md","query":{}}">
Papers
arxiv:2608.23670

Automata from Agent Traces: Failure and Next-Step Prediction

Published on Aug 24
· Submitted by
Seonglae Cho
on Aug 26
Authors:

Abstract

LLM agent traces are compressed into compact finite-state machines that enable accurate next-step and failure prediction for safety auditing and runtime monitoring.

LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.

Community

You write Agent State Machine and it drives the program.
This is the inverse: take a corpus of agent traces, recover the machine that produced them.
The machine belongs to the harness, not the LLM.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23670
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.23670 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.23670 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.23670 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers