MARS is a relay pipeline for RAG-grounded agents collaborative code generation for competitive programming. </p>\n<p>To solve the task, a team of at<br>most three agents is formed from a pool of available specialists. Each agent's turn runs code generation, public-test execution, and self-check/handoff; repair code is<br>rerun locally before the current code and relay packet move to the next specialist or final submission.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/92i27ZESEvrAbTNp2sjAn.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/92i27ZESEvrAbTNp2sjAn.png\" alt=\"relay_pipeline\"></a></p>\n","updatedAt":"2026-08-26T23:29:35.342Z","author":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","fullname":"Alsu Sagirova","name":"alsu-sagirova","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8936337232589722},"editors":["alsu-sagirova"],"editorAvatarUrls":["/avatars/1201b8282664c2d8c18beaba2396c03b.svg"],"reactions":[],"isReport":false}},{"id":"6a8f901f1a17dcac476c2528","author":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","fullname":"Alsu Sagirova","name":"alsu-sagirova","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-08-27T01:17:19.000Z","type":"comment","data":{"edited":true,"hidden":false,"latest":{"raw":"MARS reaches 0.624 pass rate at 2.3 recorded pipeline stages per task (+14.4 percentage points over direct prompting), closing most of the gap to CodeSIM (0.731) at 3.3x lower wall-clock cost and substantially smaller variance in per-task token spend.\n\n","html":"<p>MARS reaches 0.624 pass rate at 2.3 recorded pipeline stages per task (+14.4 percentage points over direct prompting), closing most of the gap to CodeSIM (0.731) at 3.3x lower wall-clock cost and substantially smaller variance in per-task token spend.<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/nWZ2x-DNl7Q8lJyK4XzsK.jpeg\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/nWZ2x-DNl7Q8lJyK4XzsK.jpeg\" alt=\"02-57-25\"></a></p>\n","updatedAt":"2026-08-27T01:20:28.479Z","author":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","fullname":"Alsu Sagirova","name":"alsu-sagirova","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":2,"identifiedLanguage":{"language":"en","probability":0.30471935868263245},"editors":["alsu-sagirova"],"editorAvatarUrls":["/avatars/1201b8282664c2d8c18beaba2396c03b.svg"],"reactions":[],"isReport":false}},{"id":"6a8f90f63d53f48748816bf5","author":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","fullname":"Alsu Sagirova","name":"alsu-sagirova","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-08-27T01:20:54.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\n\n","html":"<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/iacF49ZLP4HPHEcZzklS0.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/iacF49ZLP4HPHEcZzklS0.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-08-27T01:20:54.972Z","author":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","fullname":"Alsu Sagirova","name":"alsu-sagirova","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5895702242851257},"editors":["alsu-sagirova"],"editorAvatarUrls":["/avatars/1201b8282664c2d8c18beaba2396c03b.svg"],"reactions":[],"isReport":false}},{"id":"6a8f9139901c02a8881de652","author":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","fullname":"Alsu Sagirova","name":"alsu-sagirova","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-08-27T01:22:01.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\n\n","html":"<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/faWWfyKxQ1wrb-sw48ECW.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65c0db0fbda79a18292dfbb7/faWWfyKxQ1wrb-sw48ECW.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-08-27T01:22:01.168Z","author":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","fullname":"Alsu Sagirova","name":"alsu-sagirova","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5853187441825867},"editors":["alsu-sagirova"],"editorAvatarUrls":["/avatars/1201b8282664c2d8c18beaba2396c03b.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23918","authors":[{"_id":"6a8f74992c24e8c5fab32844","name":"Andrei Mikhailov","hidden":false},{"_id":"6a8f74992c24e8c5fab32845","name":"Mikhail Burtsev","hidden":false},{"_id":"6a8f74992c24e8c5fab32846","user":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","isPro":false,"fullname":"Alsu Sagirova","user":"alsu-sagirova","type":"user","name":"alsu-sagirova"},"name":"Alsu Sagirova","status":"claimed_verified","statusLastChangedAt":"2026-08-27T00:45:04.251Z","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"MARS: Multi-Specialist LLM Relay System for Competitive Programming","submittedOnDailyBy":{"_id":"65c0db0fbda79a18292dfbb7","avatarUrl":"/avatars/1201b8282664c2d8c18beaba2396c03b.svg","isPro":false,"fullname":"Alsu Sagirova","user":"alsu-sagirova","type":"user","name":"alsu-sagirova"},"summary":"Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribute work over generic planner, coder, and debugger roles and delegate the choice of algorithmic technique to the backbone alone. We present MARS (Multi-Agent Relay of Specialized LLMs), a prompt-only framework in which each agent is a topic specialist---dynamic programming, graphs, strings, geometry, and so on---grounded by retrieval-augmented generation over an algorithm-theory corpus. Given a problem, retrieval selects a small team of relevant specialists; a starter writes an initial C++17 solution, and each subsequent turn runs the candidate against public examples in a sandbox, lets the active specialist keep, repair, or hand off the draft, and forwards a structured packet to the next specialist. A single infrastructure-fixer pass normalizes boilerplate at the end. On the CodeContests test split with Gemma 4, MARS reaches 0.624 pm 0.006 pass rate at 2.3 recorded pipeline stages per task (+14.4 percentage points over direct prompting), closing most of the gap to CodeSIM (0.731) at 3.3{times} lower wall-clock cost and substantially smaller variance in per-task token spend. The source code is available on GitHub: https://github.com/fckand/mars.","upvotes":0,"discussionId":"6a8f74992c24e8c5fab32847","githubRepo":"https://github.com/fckand/mars","githubRepoAddedBy":"user","ai_summary":"MARS uses retrieval-augmented specialist agents for algorithmic topics to iteratively generate, test, and refine C++ solutions, improving competitive programming pass rates with lower cost.","ai_keywords":["multi-agent","retrieval-augmented generation","dynamic programming","graph algorithms","string algorithms","geometry","C++17","sandbox testing","pipeline stages","token spend"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23918.md","query":{}}">
MARS: Multi-Specialist LLM Relay System for Competitive Programming
Abstract
MARS uses retrieval-augmented specialist agents for algorithmic topics to iteratively generate, test, and refine C++ solutions, improving competitive programming pass rates with lower cost.
Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribute work over generic planner, coder, and debugger roles and delegate the choice of algorithmic technique to the backbone alone. We present MARS (Multi-Agent Relay of Specialized LLMs), a prompt-only framework in which each agent is a topic specialist---dynamic programming, graphs, strings, geometry, and so on---grounded by retrieval-augmented generation over an algorithm-theory corpus. Given a problem, retrieval selects a small team of relevant specialists; a starter writes an initial C++17 solution, and each subsequent turn runs the candidate against public examples in a sandbox, lets the active specialist keep, repair, or hand off the draft, and forwards a structured packet to the next specialist. A single infrastructure-fixer pass normalizes boilerplate at the end. On the CodeContests test split with Gemma 4, MARS reaches 0.624 pm 0.006 pass rate at 2.3 recorded pipeline stages per task (+14.4 percentage points over direct prompting), closing most of the gap to CodeSIM (0.731) at 3.3{times} lower wall-clock cost and substantially smaller variance in per-task token spend. The source code is available on GitHub: https://github.com/fckand/mars.
Community
MARS is a relay pipeline for RAG-grounded agents collaborative code generation for competitive programming.
To solve the task, a team of at
most three agents is formed from a pool of available specialists. Each agent's turn runs code generation, public-test execution, and self-check/handoff; repair code is
rerun locally before the current code and relay packet move to the next specialist or final submission.

MARS reaches 0.624 pass rate at 2.3 recorded pipeline stages per task (+14.4 percentage points over direct prompting), closing most of the gap to CodeSIM (0.731) at 3.3x lower wall-clock cost and substantially smaller variance in per-task token spend.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.23918 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.23918 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.23918 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.