Hugging Face Daily Papers · · 5 min read

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic<br>complexity of attention remains a critical bottleneck, particularly during the compute-intensive<br>prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous<br>pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic<br>prototype that is still distant from production deployment. In this paper, we present FlashPrefill<br>V2, which evolves FlashPrefill from a prototype toward practical long-context serving along<br>three dimensions. First, we introduce a mean correction term that effectively suppresses the<br>approximation error, keeping performance degradation manageable even at extreme sparsity<br>levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp<br>specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements.<br>Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing<br>integration as an attention backend in modern inference frameworks such as SGLang.</p>\n","updatedAt":"2026-08-21T04:07:50.630Z","author":{"_id":"65e7be93a5deaa480d51a88c","avatarUrl":"/avatars/44bf42dbde5ccc64114a36f1cfdad635.svg","fullname":"Qihang Fan","name":"aldjalkdf","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8778753280639648},"editors":["aldjalkdf"],"editorAvatarUrls":["/avatars/44bf42dbde5ccc64114a36f1cfdad635.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.19758","authors":[{"_id":"6a87ce7189e517cbfd75dc6d","name":"Qihang Fan","hidden":false},{"_id":"6a87ce7189e517cbfd75dc6e","name":"Huaibo Huang","hidden":false},{"_id":"6a87ce7189e517cbfd75dc6f","name":"Zhiying Wu","hidden":false},{"_id":"6a87ce7189e517cbfd75dc70","name":"Bingning Wang","hidden":false},{"_id":"6a87ce7189e517cbfd75dc71","name":"Ran He","hidden":false}],"publishedAt":"2026-08-20T08:02:55.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving","submittedOnDailyBy":{"_id":"65e7be93a5deaa480d51a88c","avatarUrl":"/avatars/44bf42dbde5ccc64114a36f1cfdad635.svg","isPro":false,"fullname":"Qihang Fan","user":"aldjalkdf","type":"user","name":"aldjalkdf"},"summary":"Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.","upvotes":4,"discussionId":"6a87ce7189e517cbfd75dc72","githubRepo":"https://github.com/qhfan/FlashPrefillv2","githubRepoAddedBy":"user","ai_summary":"FlashPrefill V2 improves long-context serving via mean-corrected sparse attention, optimized GPU operators, and framework integration, achieving large speedups over dense baselines.","ai_keywords":["FlashPrefill V2","long-context modeling","sparse attention","mean correction","PackGQA","warp specialization","pingpong pipelining","FlashAttention-3/4","FP8 inference","paged KV cache","continuous batching"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"68f4abf8f64bb4002a21a428","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/qoGuyvqteEz6gnrIxROr-.png","isPro":false,"fullname":"Boning Cui","user":"Bc-AI","type":"user"},{"_id":"68b2a4157f881fc640ba7d80","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lMTgr3pe7pOHtMe7bVF7F.png","isPro":false,"fullname":"khtsly","user":"khtsly","type":"user"},{"_id":"65e7be93a5deaa480d51a88c","avatarUrl":"/avatars/44bf42dbde5ccc64114a36f1cfdad635.svg","isPro":false,"fullname":"Qihang Fan","user":"aldjalkdf","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.19758.md","query":{}}">
Papers
arxiv:2608.19758

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Published on Aug 20
· Submitted by
Qihang Fan
on Aug 21
Authors:
,

Abstract

FlashPrefill V2 improves long-context serving via mean-corrected sparse attention, optimized GPU operators, and framework integration, achieving large speedups over dense baselines.

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.

Community

Paper submitter about 4 hours ago

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic
complexity of attention remains a critical bottleneck, particularly during the compute-intensive
prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous
pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic
prototype that is still distant from production deployment. In this paper, we present FlashPrefill
V2, which evolves FlashPrefill from a prototype toward practical long-context serving along
three dimensions. First, we introduce a mean correction term that effectively suppresses the
approximation error, keeping performance degradation manageable even at extreme sparsity
levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp
specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements.
Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing
integration as an attention backend in modern inference frameworks such as SGLang.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.19758
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.19758 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.19758 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.19758 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers