WithEveryone generates coherent group images from five to ten reference identities</p>\n","updatedAt":"2026-08-21T02:02:05.920Z","author":{"_id":"634bde123d11eaedd889e277","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1665916392312-noauth.png","fullname":"Hengyuan Xu","name":"DobyXu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.760926365852356},"editors":["DobyXu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1665916392312-noauth.png"],"reactions":[{"reaction":"🔥","users":["wchengad"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.20336","authors":[{"_id":"6a87af9a89e517cbfd75dbb9","name":"Hengyuan Xu","hidden":false},{"_id":"6a87af9a89e517cbfd75dbba","name":"Qixun Wang","hidden":false},{"_id":"6a87af9a89e517cbfd75dbbb","name":"Yiji Cheng","hidden":false},{"_id":"6a87af9a89e517cbfd75dbbc","name":"Miles Yang","hidden":false},{"_id":"6a87af9a89e517cbfd75dbbd","name":"Zhao Zhong","hidden":false},{"_id":"6a87af9a89e517cbfd75dbbe","name":"Wei Cheng","hidden":false},{"_id":"6a87af9a89e517cbfd75dbbf","name":"Xingjun Ma","hidden":false},{"_id":"6a87af9a89e517cbfd75dbc0","name":"Yu-gang Jiang","hidden":false}],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"WithEveryone: Unified Planning and Identity Grounding for Group Image Generation","submittedOnDailyBy":{"_id":"634bde123d11eaedd889e277","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1665916392312-noauth.png","isPro":false,"fullname":"Hengyuan Xu","user":"DobyXu","type":"user","name":"DobyXu"},"summary":"Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\\% of the requested identities with a duplicate rate of only 2.8\\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.","upvotes":30,"discussionId":"6a87af9a89e517cbfd75dbc1","projectPage":"https://doby-xu.github.io/WithEveryone/","githubRepo":"https://github.com/doby-xu/WithEveryone","githubRepoAddedBy":"user","ai_summary":"WithEveryone enables reliable identity-preserving group image generation for up to ten people by grounding identities to layout plans and using region-based identity losses.","ai_keywords":["identity-preserving image generation","addressed token","identity--layout plan","Layout-Grounded ID Loss","ID Representation Forcing","identity-disjoint benchmark","target-context identity similarity"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":4,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"634bde123d11eaedd889e277","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1665916392312-noauth.png","isPro":false,"fullname":"Hengyuan Xu","user":"DobyXu","type":"user"},{"_id":"68f1c5e84806540fd1aa3eb9","avatarUrl":"/avatars/1726d59a7801ab88296b55eb78a1b009.svg","isPro":false,"fullname":"Ji Yumeng","user":"guojiangqianchilang","type":"user"},{"_id":"64cb54bcbb5d195b99186e15","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64cb54bcbb5d195b99186e15/Z7ziJXjGLXN5rBEYC0sqx.png","isPro":false,"fullname":"SII-Xingjun Ma","user":"xingjunm","type":"user"},{"_id":"655101623fe6c0b1f8b58987","avatarUrl":"/avatars/4d36a4988e6011fec3ceac2b59938c3a.svg","isPro":false,"fullname":"Jiabin Hua","user":"Ammmob","type":"user"},{"_id":"68ea07d4c9dda63b9d59eca3","avatarUrl":"/avatars/e37b415735bc33afe9a31a9248591db4.svg","isPro":false,"fullname":"WithAnyone","user":"WithAnyone","type":"user"},{"_id":"67ff545a3e4487750889023d","avatarUrl":"/avatars/38538ba7576e34741cd9d20c1834c1ab.svg","isPro":false,"fullname":"Tingshu Mou","user":"Multitray","type":"user"},{"_id":"64b914c8ace99c0723ad83a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b914c8ace99c0723ad83a9/B4gxNByeVY_xaOcjwiN1j.jpeg","isPro":false,"fullname":"Wei Cheng","user":"wchengad","type":"user"},{"_id":"6685f80d6390df2eff39123f","avatarUrl":"/avatars/3d7c347c94979a7784267dd7a0c03a1d.svg","isPro":false,"fullname":"yiliucs","user":"yiliucs","type":"user"},{"_id":"6a476dbbd94f9dfc3733b8f0","avatarUrl":"/avatars/889f00d52238e19269b8d20ce96b8643.svg","isPro":false,"fullname":"tyc","user":"tyc111","type":"user"},{"_id":"63db6d6e06311e8b53e0072a","avatarUrl":"/avatars/9c9c9ca204fa4e4d6ad5d46899d26b09.svg","isPro":false,"fullname":"yangyiying","user":"yangyangyiying","type":"user"},{"_id":"6729d80b26cec81fa25edaa4","avatarUrl":"/avatars/1d6db3a3386f725443439fc86c58e4af.svg","isPro":false,"fullname":"Karamata","user":"SCFW2019","type":"user"},{"_id":"67d42f56fe3b3bee48d8f7b2","avatarUrl":"/avatars/290b0be37298baa47e94357e22f3e67e.svg","isPro":false,"fullname":"Lin","user":"hf857","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.20336.md","query":{}}">
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
Abstract
WithEveryone enables reliable identity-preserving group image generation for up to ten people by grounding identities to layout plans and using region-based identity losses.
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.
Community
WithEveryone generates coherent group images from five to ten reference identities
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.20336 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.20336 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.20336 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.