Hugging Face Daily Papers · · 4 min read

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

GS-Voxel converts pre-optimized 3D Gaussian Splatting reconstructions into sparse, structured latents without additional per-scene fitting. By combining a factorized geometry-and-attribute VAE with image-conditioned flow matching, it enables 3DGS generation and scalable large-area synthesis beyond a single training crop.<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/6433bc32a4c9c55871a42812/vM_xq1o6Ste3fopYvfC5V.jpeg\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6433bc32a4c9c55871a42812/vM_xq1o6Ste3fopYvfC5V.jpeg\" alt=\"teaser_v6\"></a></p>\n","updatedAt":"2026-08-19T05:43:51.045Z","author":{"_id":"6433bc32a4c9c55871a42812","avatarUrl":"/avatars/c27c063d02419b8ba32657514c07e3b2.svg","fullname":"qian#143","name":"qian43","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8066897988319397},"editors":["qian43"],"editorAvatarUrls":["/avatars/c27c063d02419b8ba32657514c07e3b2.svg"],"reactions":[],"isReport":false}},{"id":"6a85424f72587e599a06a073","author":{"_id":"6433bc32a4c9c55871a42812","avatarUrl":"/avatars/c27c063d02419b8ba32657514c07e3b2.svg","fullname":"qian#143","name":"qian43","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false},"createdAt":"2026-08-19T05:42:39.000Z","type":"comment","data":{"edited":true,"hidden":true,"hiddenBy":"","latest":{"raw":"This comment has been hidden","html":"This comment has been hidden","updatedAt":"2026-08-19T05:47:33.569Z","author":{"_id":"6433bc32a4c9c55871a42812","avatarUrl":"/avatars/c27c063d02419b8ba32657514c07e3b2.svg","fullname":"qian#143","name":"qian43","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"editors":[],"editorAvatarUrls":[],"reactions":[]}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17988","authors":[{"_id":"6a851499536bdd3bdd48f78d","name":"Ming Qian","hidden":false},{"_id":"6a851499536bdd3bdd48f78e","name":"Zijian Wang","hidden":false},{"_id":"6a851499536bdd3bdd48f78f","name":"Minchao Sun","hidden":false},{"_id":"6a851499536bdd3bdd48f790","name":"Jincheng Xiong","hidden":false},{"_id":"6a851499536bdd3bdd48f791","name":"Hang Zhang","hidden":false},{"_id":"6a851499536bdd3bdd48f792","name":"Mu Xu","hidden":false},{"_id":"6a851499536bdd3bdd48f793","name":"Chi Wang","hidden":false},{"_id":"6a851499536bdd3bdd48f794","name":"Baoquan Chen","hidden":false}],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation","submittedOnDailyBy":{"_id":"6433bc32a4c9c55871a42812","avatarUrl":"/avatars/c27c063d02419b8ba32657514c07e3b2.svg","isPro":false,"fullname":"qian#143","user":"qian43","type":"user","name":"qian43"},"summary":"Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.","upvotes":3,"discussionId":"6a851499536bdd3bdd48f795","ai_summary":"GS-Voxel converts unstructured 3D Gaussian reconstructions into sparse structured latents to enable scalable aerial scene generation via flow models.","ai_keywords":["3D Gaussian Splatting","3DGS","GS-Voxel","sparse active voxels","factorized VAE","sparse 3D latents","flow models","tiled inference"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"641415d08900ef6afa2fcb73","name":"acvlab","fullname":"Alibaba AMAP CV Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6414106ce7d5f817d204e160/dfveRtrRy8Xn7QpG684zl.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6433bc32a4c9c55871a42812","avatarUrl":"/avatars/c27c063d02419b8ba32657514c07e3b2.svg","isPro":false,"fullname":"qian#143","user":"qian43","type":"user"},{"_id":"690c35c8dd10faadd5ae2f80","avatarUrl":"/avatars/993d7d59f7a0b070306bb10a97cb7c35.svg","isPro":true,"fullname":"Sungwoo Park","user":"swpark5","type":"user"},{"_id":"664d930f4b870dd167473c1c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/664d930f4b870dd167473c1c/TXVEPGvkhftdI_xE1mluu.jpeg","isPro":false,"fullname":"Andy Guan","user":"andytonglove","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"641415d08900ef6afa2fcb73","name":"acvlab","fullname":"Alibaba AMAP CV Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6414106ce7d5f817d204e160/dfveRtrRy8Xn7QpG684zl.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17988.md","query":{}}">
Papers
arxiv:2608.17988

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

Published on Aug 18
· Submitted by
qian#143
on Aug 19
Authors:
,

Abstract

GS-Voxel converts unstructured 3D Gaussian reconstructions into sparse structured latents to enable scalable aerial scene generation via flow models.

Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.

Community

GS-Voxel converts pre-optimized 3D Gaussian Splatting reconstructions into sparse, structured latents without additional per-scene fitting. By combining a factorized geometry-and-attribute VAE with image-conditioned flow matching, it enables 3DGS generation and scalable large-area synthesis beyond a single training crop.
teaser_v6

Paper submitter about 2 hours ago
This comment has been hidden
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.17988
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.17988 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.17988 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.17988 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers