Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at 256×256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512×512 resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at <a href=\"https://github.com/ysng123/EG-FM\" rel=\"nofollow\">https://github.com/ysng123/EG-FM</a>.</p>\n","updatedAt":"2026-08-19T03:10:08.879Z","author":{"_id":"68c156ad4e2abd9edcfa0953","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2BtrMBkobql_y0ixboyg9.png","fullname":"g","name":"ysng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8847761750221252},"editors":["ysng"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2BtrMBkobql_y0ixboyg9.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05811","authors":[{"_id":"6a758076e1228e04b3238320","user":{"_id":"68c156ad4e2abd9edcfa0953","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2BtrMBkobql_y0ixboyg9.png","isPro":false,"fullname":"g","user":"ysng","type":"user","name":"ysng"},"name":"Haoyang Tong","status":"admin_assigned","statusLastChangedAt":"2026-08-18T17:18:33.908Z","hidden":false},{"_id":"6a758076e1228e04b3238321","name":"Yu He","hidden":false},{"_id":"6a758076e1228e04b3238322","name":"Fang Li","hidden":false},{"_id":"6a758076e1228e04b3238323","name":"Lichen Ma","hidden":false},{"_id":"6a758076e1228e04b3238324","name":"Jingling Fu","hidden":false},{"_id":"6a758076e1228e04b3238325","name":"Dong Chen","hidden":false},{"_id":"6a758076e1228e04b3238326","name":"Zhen Chen","hidden":false},{"_id":"6a758076e1228e04b3238327","name":"Junshi Huang","hidden":false},{"_id":"6a758076e1228e04b3238328","name":"Jie Cao","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"Energy-Guided Flow Matching","submittedOnDailyBy":{"_id":"68c156ad4e2abd9edcfa0953","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2BtrMBkobql_y0ixboyg9.png","isPro":false,"fullname":"g","user":"ysng","type":"user","name":"ysng"},"summary":"Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at 256 times 256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512 times 512 resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at https://github.com/ysng123/EG-FM.","upvotes":6,"discussionId":"6a758076e1228e04b3238329","projectPage":"https://ysng123.github.io/EG-FM/","githubRepo":"https://github.com/ysng123/EG-FM","githubRepoAddedBy":"user","ai_summary":"Energy-Guided Flow Matching improves generative quality by progressively revealing high-frequency details through a moving endpoint and adaptive scheduling, reducing training cost and achieving state-of-the-art FID scores.","ai_keywords":["flow matching","heat-kernel-filtered endpoint","energy-guided scheduling","velocity retargeting","class-conditional image generation","text-to-image generation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":15},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66d0407dc652abb652610e6d","avatarUrl":"/avatars/929098ad55d73fd9191a4efd0dff2155.svg","isPro":false,"fullname":"sng","user":"ysd123321","type":"user"},{"_id":"68c156ad4e2abd9edcfa0953","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2BtrMBkobql_y0ixboyg9.png","isPro":false,"fullname":"g","user":"ysng","type":"user"},{"_id":"630f612fcc8ed75decb4796e","avatarUrl":"/avatars/261f03bb0c926d66993df9560abb74fc.svg","isPro":true,"fullname":"Lucas","user":"xxlucas","type":"user"},{"_id":"67bbade8a8c89b98ec377944","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67bbade8a8c89b98ec377944/HPtKDo8fnKr4OxpN1Z17D.png","isPro":false,"fullname":"Urodoc Oncall","user":"UDCAI","type":"user"},{"_id":"64ac26b23215a18926f2bb28","avatarUrl":"/avatars/022865242846a294488530e3e1ebd303.svg","isPro":true,"fullname":"Fang Li","user":"Neesky","type":"user"},{"_id":"66fb03d6b505f1a04c39d935","avatarUrl":"/avatars/e9b830c460ec02037758c9b3469bb8ad.svg","isPro":false,"fullname":"Xuanlang Dai","user":"XuanlangDai","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05811.md","query":{}}">
Energy-Guided Flow Matching
Published on Aug 7
· Submitted by g on Aug 19 Abstract
Energy-Guided Flow Matching improves generative quality by progressively revealing high-frequency details through a moving endpoint and adaptive scheduling, reducing training cost and achieving state-of-the-art FID scores.
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at 256 times 256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512 times 512 resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at https://github.com/ysng123/EG-FM.
Community
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at 256×256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512×512 resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at https://github.com/ysng123/EG-FM.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.05811 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.05811 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.