Video Generative Models as Geometry Learner
Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.
Abstract
GeoNeXt repurposes pretrained video generative models as a unified framework for geometry estimation via next-frame prediction, enabling efficient joint modeling of depth and surface normals with minimal labeled data.
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.
Models citing this paper
No model linking this paper
Datasets citing this paper
No dataset linking this paper
Spaces citing this paper
No Space linking this paper
Collections including this paper
No Collection including this paper
More from Hugging Face Daily Papers
-
Rubric-to-Code Credit Assignment for Reinforcement Learning
Aug 31
-
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Aug 31
-
Fast Weight Attention for Continual Learning
Aug 31
-
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
Aug 31
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.