We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs)<br>to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional<br>correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs’ ability to exploit evolving GPU architectures.</p>\n","updatedAt":"2026-08-19T15:29:54.930Z","author":{"_id":"65a76ff1e504d9738d636217","avatarUrl":"/avatars/26bf5e3f19057835ee95d72c24904d77.svg","fullname":"Genghan Zhang","name":"Genghan","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8714826703071594},"editors":["Genghan"],"editorAvatarUrls":["/avatars/26bf5e3f19057835ee95d72c24904d77.svg"],"reactions":[],"isReport":false}},{"id":"6a85d20359d6a9a5cf9f885e","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-19T15:55:47.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?](https://huggingface.co/papers/2608.04450) (2026)\n* [Harness Engineering for LLM-Driven GPU Kernel Generation](https://huggingface.co/papers/2607.17979) (2026)\n* [Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent](https://huggingface.co/papers/2607.14541) (2026)\n* [KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?](https://huggingface.co/papers/2607.16241) (2026)\n* [CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution](https://huggingface.co/papers/2608.12629) (2026)\n* [KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation](https://huggingface.co/papers/2607.27231) (2026)\n* [Rethinking Agentic Kernel Generation for Emerging Accelerators](https://huggingface.co/papers/2608.00894) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2608.04450\">CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.17979\">Harness Engineering for LLM-Driven GPU Kernel Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.14541\">Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.16241\">KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.12629\">CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.27231\">KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.00894\">Rethinking Agentic Kernel Generation for Emerging Accelerators</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-19T15:55:47.233Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7072325944900513},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17379","authors":[{"_id":"6a8534f2536bdd3bdd48f885","user":{"_id":"65a76ff1e504d9738d636217","avatarUrl":"/avatars/26bf5e3f19057835ee95d72c24904d77.svg","isPro":true,"fullname":"Genghan Zhang","user":"Genghan","type":"user","name":"Genghan"},"name":"Genghan Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-19T08:45:04.462Z","hidden":false},{"_id":"6a8534f2536bdd3bdd48f886","name":"Yixin Dong","hidden":false},{"_id":"6a8534f2536bdd3bdd48f887","name":"Chengze Fan","hidden":false},{"_id":"6a8534f2536bdd3bdd48f888","name":"Zhichen Zeng","hidden":false},{"_id":"6a8534f2536bdd3bdd48f889","name":"Yueming Yuan","hidden":false},{"_id":"6a8534f2536bdd3bdd48f88a","name":"Shaowei Zhu","hidden":false},{"_id":"6a8534f2536bdd3bdd48f88b","name":"Kunle Olukotun","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65a76ff1e504d9738d636217/cbYJVfn8UbJGHsiFBI-_e.png"],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX","submittedOnDailyBy":{"_id":"65a76ff1e504d9738d636217","avatarUrl":"/avatars/26bf5e3f19057835ee95d72c24904d77.svg","isPro":true,"fullname":"Genghan Zhang","user":"Genghan","type":"user","name":"Genghan"},"summary":"We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.","upvotes":1,"discussionId":"6a8534f2536bdd3bdd48f88c","projectPage":"https://zhang677.github.io/blog_md/ptxbench.html","githubRepo":"https://github.com/zhang677/PTXBench","githubRepoAddedBy":"user","ai_summary":"PTXBench evaluates large language models on architecture-specific GPU kernel optimization, revealing uneven success and performance gaps that supervised fine-tuning only partially addresses.","ai_keywords":["PTXBench","PTX","GPU kernel optimization","GEMM","attention workloads","H100","B200","supervised fine-tuning","repair-conditioned training","reasoning teacher"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"672c672dcf09d152f4da04c4","name":"StanfordUniversity","fullname":"Stanford University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/vJI0POlzGMXL2878t1vz2.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"672c672dcf09d152f4da04c4","name":"StanfordUniversity","fullname":"Stanford University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/vJI0POlzGMXL2878t1vz2.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17379.md","query":{}}">
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
Abstract
PTXBench evaluates large language models on architecture-specific GPU kernel optimization, revealing uneven success and performance gaps that supervised fine-tuning only partially addresses.
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.
Community
This comment has been hidden (marked as Resolved) We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs)
to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional
correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs’ ability to exploit evolving GPU architectures.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.17379 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.17379 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.17379 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.