Hugging Face Daily Papers · · 4 min read

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🚀 Can one vision encoder understand, reconstruct, and generate?</p>\n<p>UniSpace explores this question by reparameterizing the patch embedding of pretrained ViTs. Instead of adding separate semantic and reconstruction encoders, it keeps the original semantic pathway and introduces a reconstruction-aware pathway to recover fine-grained visual details.</p>\n<p>The key insight is fascinating: frozen ViT blocks may already contain rich visual information — the bottleneck is how we parameterize the input tokens.</p>\n<p>A step towards more unified visual representations for multimodal models!</p>\n","updatedAt":"2026-08-24T09:21:47.638Z","author":{"_id":"66ab48a150bd6711f3b60ca0","avatarUrl":"/avatars/a526ea346d1f78535da1b60567670465.svg","fullname":"yjb","name":"yjb6","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7631049156188965},"editors":["yjb6"],"editorAvatarUrls":["/avatars/a526ea346d1f78535da1b60567670465.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.08676","authors":[{"_id":"6a7ab6cc019ce76dc7b3ab13","user":{"_id":"66ab48a150bd6711f3b60ca0","avatarUrl":"/avatars/a526ea346d1f78535da1b60567670465.svg","isPro":false,"fullname":"yjb","user":"yjb6","type":"user","name":"yjb6"},"name":"Jinbo Yan","status":"claimed_verified","statusLastChangedAt":"2026-08-21T13:14:14.239Z","hidden":false},{"_id":"6a7ab6cc019ce76dc7b3ab14","name":"Limeng Qiao","hidden":false},{"_id":"6a7ab6cc019ce76dc7b3ab15","name":"Jie Qin","hidden":false},{"_id":"6a7ab6cc019ce76dc7b3ab16","name":"Junyan He","hidden":false},{"_id":"6a7ab6cc019ce76dc7b3ab17","name":"Feize Wu","hidden":false},{"_id":"6a7ab6cc019ce76dc7b3ab18","name":"Guanglu Wan","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66ab48a150bd6711f3b60ca0/ZZajIjVyH22Pxj_wEyP3F.png"],"publishedAt":"2026-08-09T00:00:00.000Z","submittedOnDailyAt":"2026-08-24T00:00:00.000Z","title":"UniSpace: Unified Visual Representation and Scalable Multimodal Modeling","submittedOnDailyBy":{"_id":"66ab48a150bd6711f3b60ca0","avatarUrl":"/avatars/a526ea346d1f78535da1b60567670465.svg","isPro":false,"fullname":"yjb","user":"yjb6","type":"user","name":"yjb6"},"summary":"Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and editing can be modeled in a single visual representation space built from a pretrained semantic ViT. We show that the frozen Transformer blocks of a semantic ViT are not intrinsically unable to preserve visual details. Instead, the original patch parameterization drives the representation toward semantic abstraction, making fine-grained information difficult to recover from the final tokens. Based on this observation, we introduce Patch Reparameterization, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks. The resulting unified representation preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off. We further scale this representation into UniSpace, an 8B Mixture-of-Transformer-Experts model that performs understanding, generation, and editing in the same visual space without a separate VAE pathway. System-level evaluations demonstrate practical text-to-image generation and instruction-based image editing, showing that a reparameterized pretrained ViT can serve as a unified visual interface for scalable multimodal modeling.","upvotes":7,"discussionId":"6a7ab6cc019ce76dc7b3ab19","projectPage":"https://yjb6.github.io/UniSpace/","githubRepo":"https://github.com/yjb6/UniSpace","githubRepoAddedBy":"user","ai_summary":"A reparameterized pretrained vision transformer unifies semantic understanding, high-fidelity reconstruction, and image generation within a single visual space without requiring a separate VAE.","ai_keywords":["semantic ViT","Patch Reparameterization","reconstruction-aware patch embedding","Mixture-of-Transformer-Experts","UniSpace","unified visual interface"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"68b28d79a176a9beb30d2049","name":"meituan-longcat","fullname":"LongCat","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68a2a29ab9d4c5698e02c747/CDCAx7X7rXDt7xjI-DoxG.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66ab48a150bd6711f3b60ca0","avatarUrl":"/avatars/a526ea346d1f78535da1b60567670465.svg","isPro":false,"fullname":"yjb","user":"yjb6","type":"user"},{"_id":"65af7c59bff3ae9f33990ed5","avatarUrl":"/avatars/3de08df23883e1a0b7f36d2a991f9f24.svg","isPro":false,"fullname":"FeizeWu","user":"FeizeWu","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"68f59ae49315a06ad9a01464","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68f59ae49315a06ad9a01464/welpgtUr6TY9Qa7nPBigg.jpeg","isPro":false,"fullname":"Sean Yu","user":"yushaohan","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"69a5cba5ee290d6bb49457b8","avatarUrl":"/avatars/f80c17c13d6baf6bcd375d31efe21116.svg","isPro":false,"fullname":"Darrow O'Lykos","user":"darrowoflykos","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68b28d79a176a9beb30d2049","name":"meituan-longcat","fullname":"LongCat","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68a2a29ab9d4c5698e02c747/CDCAx7X7rXDt7xjI-DoxG.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.08676.md","query":{}}">
Papers
arxiv:2608.08676

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

Published on Aug 9
· Submitted by
yjb
on Aug 24
Authors:

Abstract

A reparameterized pretrained vision transformer unifies semantic understanding, high-fidelity reconstruction, and image generation within a single visual space without requiring a separate VAE.

Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and editing can be modeled in a single visual representation space built from a pretrained semantic ViT. We show that the frozen Transformer blocks of a semantic ViT are not intrinsically unable to preserve visual details. Instead, the original patch parameterization drives the representation toward semantic abstraction, making fine-grained information difficult to recover from the final tokens. Based on this observation, we introduce Patch Reparameterization, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks. The resulting unified representation preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off. We further scale this representation into UniSpace, an 8B Mixture-of-Transformer-Experts model that performs understanding, generation, and editing in the same visual space without a separate VAE pathway. System-level evaluations demonstrate practical text-to-image generation and instruction-based image editing, showing that a reparameterized pretrained ViT can serve as a unified visual interface for scalable multimodal modeling.

Community

Paper author Paper submitter about 8 hours ago

🚀 Can one vision encoder understand, reconstruct, and generate?

UniSpace explores this question by reparameterizing the patch embedding of pretrained ViTs. Instead of adding separate semantic and reconstruction encoders, it keeps the original semantic pathway and introduces a reconstruction-aware pathway to recover fine-grained visual details.

The key insight is fascinating: frozen ViT blocks may already contain rich visual information — the bottleneck is how we parameterize the input tokens.

A step towards more unified visual representations for multimodal models!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.08676
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.08676 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers