Hugging Face Daily Papers · · 7 min read

Luce: Relightable Gaussians for 3D Asset Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

High-fidelity image-to-3D generation requires a 3D representation that captures<br>both geometry and appearance. To support relighting and integration into standard<br>rendering pipelines, the representation should include physically based rendering<br>(PBR) modalities such as albedo, metallic-roughness, and surface normals. We<br>propose Luce, a 3D representation that unifies geometry and PBR materials within<br>a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for<br>each modality. A variational autoencoder compresses this representation into a<br>unified material-aware latent space. A rectified-flow transformer generates this la-<br>tent from a single image, conditioned on multi-layer features from a pretrained im-<br>age encoder that preserve both semantic context and fine spatial detail. The latent<br>then decodes into relightable PBR Gaussians and an optional textured mesh with<br>a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-<br>image-to-3D generation, improving FID by 28% over the strongest baseline. We<br>further introduce a benchmark of AI-generated images, on which Luce improves<br>the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce<br>generates relightable, geometrically accurate, and materially faithful assets that<br>preserve fine details such as text, logos, and inscriptions.</p>\n","updatedAt":"2026-08-29T00:15:37.086Z","author":{"_id":"63f4edf571a5d395c71d01cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f4edf571a5d395c71d01cf/yHEMEyYO0-wO3ZmI89OkY.jpeg","fullname":"Aditya Ganeshan","name":"bardofcodes","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8451366424560547},"editors":["bardofcodes"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/63f4edf571a5d395c71d01cf/yHEMEyYO0-wO3ZmI89OkY.jpeg"],"reactions":[],"isReport":false}},{"id":"6a9232e4ca63df32033b104d","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-29T01:16:20.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [InvSplat: Inverse Feed-Forward Scene Splatting](https://huggingface.co/papers/2607.02301) (2026)\n* [ExMesh++: From Multi-View Images to Relightable UV-PBR Mesh Assets via Topology-Adaptive Reconstruction and Decomposition](https://huggingface.co/papers/2608.24109) (2026)\n* [LumiTokens: 3D Relighting via Token-Space Lighting Transformation](https://huggingface.co/papers/2608.18215) (2026)\n* [DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion](https://huggingface.co/papers/2608.20759) (2026)\n* [PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation](https://huggingface.co/papers/2607.01803) (2026)\n* [LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting](https://huggingface.co/papers/2607.08016) (2026)\n* [InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\\deg} Image](https://huggingface.co/papers/2607.03990) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.02301\">InvSplat: Inverse Feed-Forward Scene Splatting</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.24109\">ExMesh++: From Multi-View Images to Relightable UV-PBR Mesh Assets via Topology-Adaptive Reconstruction and Decomposition</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.18215\">LumiTokens: 3D Relighting via Token-Space Lighting Transformation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.20759\">DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.01803\">PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.08016\">LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.03990\">InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\\deg} Image</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-29T01:16:20.833Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6639047861099243},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23943","authors":[{"_id":"6a91155ca64059bab69c36da","user":{"_id":"69432676c2970b786ef70b64","avatarUrl":"/avatars/70b94d156a91eb34ebd064189dcd4c8c.svg","isPro":false,"fullname":"Mayank Singh","user":"mayanksinghkgp","type":"user","name":"mayanksinghkgp"},"name":"Mayank Singh","status":"claimed_verified","statusLastChangedAt":"2026-08-29T00:45:04.680Z","hidden":false},{"_id":"6a91155ca64059bab69c36db","name":"Michele Stoppa","hidden":false},{"_id":"6a91155ca64059bab69c36dc","name":"Alvise Memo","hidden":false},{"_id":"6a91155ca64059bab69c36dd","name":"Rui Yu","hidden":false},{"_id":"6a91155ca64059bab69c36de","name":"Harsha Kalli","hidden":false},{"_id":"6a91155ca64059bab69c36df","name":"Srimanth Gunturi","hidden":false},{"_id":"6a91155ca64059bab69c36e0","name":"Muhammad Ahmed Riaz","hidden":false},{"_id":"6a91155ca64059bab69c36e1","name":"Behrooz Shahsavari","hidden":false},{"_id":"6a91155ca64059bab69c36e2","name":"Waleed Abdulla","hidden":false},{"_id":"6a91155ca64059bab69c36e3","name":"David E. Jacobs","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-28T00:00:00.000Z","title":"Luce: Relightable Gaussians for 3D Asset Generation","submittedOnDailyBy":{"_id":"63f4edf571a5d395c71d01cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f4edf571a5d395c71d01cf/yHEMEyYO0-wO3ZmI89OkY.jpeg","isPro":false,"fullname":"Aditya Ganeshan","user":"bardofcodes","type":"user","name":"bardofcodes"},"summary":"High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include physically based rendering (PBR) modalities such as albedo, metallic-roughness, and surface normals. We propose Luce, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for each modality. A variational autoencoder compresses this representation into a unified material-aware latent space. A rectified-flow transformer generates this latent from a single image, conditioned on multi-layer features from a pretrained image encoder that preserve both semantic context and fine spatial detail. The latent then decodes into relightable PBR Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-image-to-3D generation, improving FID by 28% over the strongest baseline. We further introduce a benchmark of AI-generated images, on which Luce improves the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce generates relightable, geometrically accurate, and materially faithful assets that preserve fine details such as text, logos, and inscriptions.","upvotes":6,"discussionId":"6a91155ca64059bab69c36e4","ai_summary":"Luce unifies geometry and PBR materials in a voxelized Gaussian cloud, using a variational autoencoder and rectified-flow transformer to generate relightable 3D assets from single images.","ai_keywords":["voxelized multimodal Gaussian cloud","PBR materials","variational autoencoder","rectified-flow transformer","multi-layer features","tangent-space normal map"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"628cbd99ef14f971b69948ab","name":"apple","fullname":"Apple","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1653390727490-5dd96eb166059660ed1ee413.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69432676c2970b786ef70b64","avatarUrl":"/avatars/70b94d156a91eb34ebd064189dcd4c8c.svg","isPro":false,"fullname":"Mayank Singh","user":"mayanksinghkgp","type":"user"},{"_id":"62f6a894c3372328414c7021","avatarUrl":"/avatars/823a48cee3ae1ee64d693d50fa74eb70.svg","isPro":false,"fullname":"Nupur Kumari","user":"nupurkmr9","type":"user"},{"_id":"65b594f4188d9466f3afaa00","avatarUrl":"/avatars/d04ac8fb29d67207563ab1aec4cc5355.svg","isPro":true,"fullname":"Ben Sha","user":"bensh","type":"user"},{"_id":"6a56af93fe5a94ee3b649acd","avatarUrl":"/avatars/864facb0c9d29b2f09c88e085ba75cba.svg","isPro":false,"fullname":"ise","user":"Alv9898","type":"user"},{"_id":"6a529655066028dd8a66ec52","avatarUrl":"/avatars/a2d92ebb13bb102fe59309f4db2c3e36.svg","isPro":false,"fullname":"Barkha Rani","user":"ranibarkha","type":"user"},{"_id":"63f4edf571a5d395c71d01cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f4edf571a5d395c71d01cf/yHEMEyYO0-wO3ZmI89OkY.jpeg","isPro":false,"fullname":"Aditya Ganeshan","user":"bardofcodes","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"628cbd99ef14f971b69948ab","name":"apple","fullname":"Apple","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1653390727490-5dd96eb166059660ed1ee413.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23943.md","query":{}}">
Papers
arxiv:2608.23943

Luce: Relightable Gaussians for 3D Asset Generation

Published on Aug 25
· Submitted by
Aditya Ganeshan
on Aug 28
Authors:

Abstract

Luce unifies geometry and PBR materials in a voxelized Gaussian cloud, using a variational autoencoder and rectified-flow transformer to generate relightable 3D assets from single images.

High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include physically based rendering (PBR) modalities such as albedo, metallic-roughness, and surface normals. We propose Luce, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for each modality. A variational autoencoder compresses this representation into a unified material-aware latent space. A rectified-flow transformer generates this latent from a single image, conditioned on multi-layer features from a pretrained image encoder that preserve both semantic context and fine spatial detail. The latent then decodes into relightable PBR Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-image-to-3D generation, improving FID by 28% over the strongest baseline. We further introduce a benchmark of AI-generated images, on which Luce improves the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce generates relightable, geometrically accurate, and materially faithful assets that preserve fine details such as text, logos, and inscriptions.

Community

Paper submitter about 2 hours ago

High-fidelity image-to-3D generation requires a 3D representation that captures
both geometry and appearance. To support relighting and integration into standard
rendering pipelines, the representation should include physically based rendering
(PBR) modalities such as albedo, metallic-roughness, and surface normals. We
propose Luce, a 3D representation that unifies geometry and PBR materials within
a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for
each modality. A variational autoencoder compresses this representation into a
unified material-aware latent space. A rectified-flow transformer generates this la-
tent from a single image, conditioned on multi-layer features from a pretrained im-
age encoder that preserve both semantic context and fine spatial detail. The latent
then decodes into relightable PBR Gaussians and an optional textured mesh with
a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-
image-to-3D generation, improving FID by 28% over the strongest baseline. We
further introduce a benchmark of AI-generated images, on which Luce improves
the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce
generates relightable, geometrically accurate, and materially faithful assets that
preserve fine details such as text, logos, and inscriptions.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23943
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.23943 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.23943 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.23943 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers