You write Agent State Machine and it drives the program.<br>This is the inverse: take a corpus of agent traces, recover the machine that produced them.<br>The machine belongs to the harness, not the LLM.</p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/64105805928400b416439f10/CAaswF8FGlReDAooYT-t7.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-08-26T23:19:24.247Z","author":{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","fullname":"Seonglae Cho","name":"seonglae","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.9164570569992065},"editors":["seonglae"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23670","authors":[{"_id":"6a8efd672c24e8c5fab326ac","user":{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae","type":"user","name":"seonglae"},"name":"Seonglae Cho","status":"claimed_verified","statusLastChangedAt":"2026-08-26T16:45:04.654Z","hidden":false},{"_id":"6a8efd672c24e8c5fab326ad","name":"Franklin Cardenoso Fernandez","hidden":false},{"_id":"6a8efd672c24e8c5fab326ae","name":"Umar Mohammed","hidden":false},{"_id":"6a8efd672c24e8c5fab326af","name":"Zekun Wu","hidden":false},{"_id":"6a8efd672c24e8c5fab326b0","name":"Kleyton Da Costa","hidden":false},{"_id":"6a8efd672c24e8c5fab326b1","name":"Ilham Wicaksono","hidden":false},{"_id":"6a8efd672c24e8c5fab326b2","name":"Adriano Koshiyama","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"Automata from Agent Traces: Failure and Next-Step Prediction","submittedOnDailyBy":{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae","type":"user","name":"seonglae"},"summary":"LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.","upvotes":4,"discussionId":"6a8efd672c24e8c5fab326b3","projectPage":"https://seongland.com/article/asg","ai_summary":"LLM agent traces are compressed into compact finite-state machines that enable accurate next-step and failure prediction for safety auditing and runtime monitoring.","ai_keywords":["finite-state machine","LLM agents","behavioral topology","next-step prediction","failure prediction","runtime monitoring","safety auditing","cross-run topology"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"64a32e214c758629715d2607","name":"holistic-ai","fullname":"Holistic AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648c6ae855101b9f9affdbc7/MFj_bUiWw-TFpouPKQian.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64105805928400b416439f10","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64105805928400b416439f10/i0jFLo47RTDeNl26hiR9y.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae","type":"user"},{"_id":"68383cfda2d6f83cf1ee88a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/mGNmuxzz1zaPZyMQFCde6.jpeg","isPro":false,"fullname":"Seonglae Cho","user":"seonglae-holistic","type":"user"},{"_id":"631e14ac473a6825f285e89d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/631e14ac473a6825f285e89d/K-6QnoeGLg8XFvbTMMdqA.jpeg","isPro":false,"fullname":"Yury Panikov","user":"panikov","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64a32e214c758629715d2607","name":"holistic-ai","fullname":"Holistic AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648c6ae855101b9f9affdbc7/MFj_bUiWw-TFpouPKQian.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23670.md","query":{}}">
Automata from Agent Traces: Failure and Next-Step Prediction
Abstract
LLM agent traces are compressed into compact finite-state machines that enable accurate next-step and failure prediction for safety auditing and runtime monitoring.
LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.
Community
You write Agent State Machine and it drives the program.
This is the inverse: take a corpus of agent traces, recover the machine that produced them.
The machine belongs to the harness, not the LLM.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.23670 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.23670 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.23670 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.