We study what happens when coding agents switch models mid-trajectory. Across Claude and GPT model pairs, escalation suffers a substantial “handoff tax,” while downshifting offers a more favorable cost–quality trade-off—and the best handoff interface depends on the switching direction.</p>\n","updatedAt":"2026-08-27T05:53:51.866Z","author":{"_id":"62f0cb39671fc964b5063aba","avatarUrl":"/avatars/02404cd78c395331b42500dd6ced35eb.svg","fullname":"Roy Ganz","name":"proy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8357841372489929},"editors":["proy"],"editorAvatarUrls":["/avatars/02404cd78c395331b42500dd6ced35eb.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24358","authors":[{"_id":"6a8fd0583bd48bb654ea690f","name":"Roy Ganz","hidden":false},{"_id":"6a8fd0583bd48bb654ea6910","name":"Mor Shpigel Nacson","hidden":false},{"_id":"6a8fd0583bd48bb654ea6911","name":"Adi Kalyanpur","hidden":false},{"_id":"6a8fd0583bd48bb654ea6912","name":"Ron Litman","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents","submittedOnDailyBy":{"_id":"62f0cb39671fc964b5063aba","avatarUrl":"/avatars/02404cd78c395331b42500dd6ced35eb.svg","isPro":false,"fullname":"Roy Ganz","user":"proy","type":"user","name":"proy"},"summary":"Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once the hard reasoning is complete. Each switch requires the receiver to continue a non-native trajectory produced by another model. We study how this handoff affects quality and cost, and how varying the trajectory information inherited by the receiver changes the outcome. Using pairs of low-cost, low-capability (LC) and high-cost, high-capability (HC) models from the Claude and GPT families, we vary handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state. Across both model families, full-trajectory escalation recovers less than half of the LC-to-HC quality gap while incurring a substantial cost premium. We term this cost-quality penalty the handoff tax. By contrast, downshift offers a favorable cost-quality point. Interestingly, the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.","upvotes":4,"discussionId":"6a8fd0593bd48bb654ea6913"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62f0cb39671fc964b5063aba","avatarUrl":"/avatars/02404cd78c395331b42500dd6ced35eb.svg","isPro":false,"fullname":"Roy Ganz","user":"proy","type":"user"},{"_id":"65e70aa92db014a1b278a857","avatarUrl":"/avatars/2f586a73408de49af10c1efe6abcb9cf.svg","isPro":false,"fullname":"Ron Litman","user":"ronlitman","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"67b4079145dc598e0f110530","avatarUrl":"/avatars/30647b6d748a4ff2d1cb1c19daacd005.svg","isPro":false,"fullname":"Harry Lu","user":"HelloWorld9724","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24358.md","query":{}}">
The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
Abstract
Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once the hard reasoning is complete. Each switch requires the receiver to continue a non-native trajectory produced by another model. We study how this handoff affects quality and cost, and how varying the trajectory information inherited by the receiver changes the outcome. Using pairs of low-cost, low-capability (LC) and high-cost, high-capability (HC) models from the Claude and GPT families, we vary handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state. Across both model families, full-trajectory escalation recovers less than half of the LC-to-HC quality gap while incurring a substantial cost premium. We term this cost-quality penalty the handoff tax. By contrast, downshift offers a favorable cost-quality point. Interestingly, the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.
Community
We study what happens when coding agents switch models mid-trajectory. Across Claude and GPT model pairs, escalation suffers a substantial “handoff tax,” while downshifting offers a more favorable cost–quality trade-off—and the best handoff interface depends on the switching direction.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.24358 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.24358 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.24358 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.