CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing</p>\n<p>Project Page: <a href=\"https://coinve200k.github.io/\" rel=\"nofollow\">https://coinve200k.github.io/</a><br>Code: <a href=\"https://github.com/coinve200k/CoinVE-200K\" rel=\"nofollow\">https://github.com/coinve200k/CoinVE-200K</a><br>Dataset: <a href=\"https://huggingface.co/datasets/FireCRT/CoinVE-200K\">https://huggingface.co/datasets/FireCRT/CoinVE-200K</a><br>Model: <a href=\"https://huggingface.co/FireCRT/CoinVE-Edit\">https://huggingface.co/FireCRT/CoinVE-Edit</a><br>Bench: <a href=\"https://huggingface.co/datasets/FireCRT/CoinVE-Bench\">https://huggingface.co/datasets/FireCRT/CoinVE-Bench</a></p>\n","updatedAt":"2026-08-19T06:01:23.225Z","author":{"_id":"6449f2dfeb7db8f70fb990f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png","fullname":"Fuchen","name":"FireCRT","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6024860739707947},"editors":["FireCRT"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17566","authors":[{"_id":"6a851536536bdd3bdd48f7d0","name":"Fuchen Long","hidden":false},{"_id":"6a851536536bdd3bdd48f7d1","name":"Cong Wang","hidden":false},{"_id":"6a851536536bdd3bdd48f7d2","name":"Zitao Gao","hidden":false},{"_id":"6a851536536bdd3bdd48f7d3","name":"Wenhao Zhong","hidden":false},{"_id":"6a851536536bdd3bdd48f7d4","name":"Yu Cheng","hidden":false},{"_id":"6a851536536bdd3bdd48f7d5","name":"Xiaolu Hou","hidden":false},{"_id":"6a851536536bdd3bdd48f7d6","name":"Yan Li","hidden":false},{"_id":"6a851536536bdd3bdd48f7d7","name":"Xiao Cao","hidden":false},{"_id":"6a851536536bdd3bdd48f7d8","name":"Xinlong Sun","hidden":false},{"_id":"6a851536536bdd3bdd48f7d9","name":"Xi Chen","hidden":false},{"_id":"6a851536536bdd3bdd48f7da","name":"Yu Liu","hidden":false}],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing","submittedOnDailyBy":{"_id":"6449f2dfeb7db8f70fb990f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png","isPro":false,"fullname":"Fuchen","user":"FireCRT","type":"user","name":"FireCRT"},"summary":"The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.","upvotes":10,"discussionId":"6a851536536bdd3bdd48f7db","projectPage":"https://coinve200k.github.io/","githubRepo":"https://github.com/coinve200k/CoinVE-200K","githubRepoAddedBy":"user","ai_summary":"A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence.","ai_keywords":["compositional instruction-guided video editing","CoinVE-200K","CoinVE-Bench","CoinVE-Edit","region-aware attention","Wan2.1-T2V-14B","Qwen3-VL-8B-Instruct","temporal consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6449f2dfeb7db8f70fb990f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png","isPro":false,"fullname":"Fuchen","user":"FireCRT","type":"user"},{"_id":"6556ea0e0f4493529783e7a4","avatarUrl":"/avatars/df404cae1511414ce648469e9f0f0714.svg","isPro":false,"fullname":"PyBigStar","user":"PyBigStar","type":"user"},{"_id":"644b96709ca7117968e75997","avatarUrl":"/avatars/3ba8fbde30feee7f117ec015ce13dc15.svg","isPro":false,"fullname":"BiBiKo_219","user":"BiBiKo","type":"user"},{"_id":"646339704ad7e61e51db904c","avatarUrl":"/avatars/a073ac880bcc3ec5f3a832da4a94ed8e.svg","isPro":false,"fullname":"SkyScraper","user":"Cong723","type":"user"},{"_id":"68834a9283873f346bfc88bc","avatarUrl":"/avatars/4643d7931857c9c3bf9b192ca8d15a65.svg","isPro":false,"fullname":"binyuanhuang","user":"foralex","type":"user"},{"_id":"6878987600ce33ca49b91d3f","avatarUrl":"/avatars/e30e08c2f8bfdb4f00103327dd43c21a.svg","isPro":false,"fullname":"cy","user":"wuucy","type":"user"},{"_id":"6496f5754a3c31df8e3139f6","avatarUrl":"/avatars/cf789d1986f976373c82b2976df4542a.svg","isPro":false,"fullname":"Zhongwei Zhang","user":"zzwustc","type":"user"},{"_id":"686349186139862d1f7d1be3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/YTthC2OhhUAu5dQ46EjlH.png","isPro":false,"fullname":"wehaozhong","user":"wenhaozhong","type":"user"},{"_id":"65f1642abf8631e7dae1db4e","avatarUrl":"/avatars/4d6256b94abeb2ff6894b1f9a2c11db7.svg","isPro":false,"fullname":"yucheng","user":"Yumic","type":"user"},{"_id":"6678dc9fa845e4470fab99fc","avatarUrl":"/avatars/c10f1ec6be0eedfb9cf07a6d2cb700a3.svg","isPro":false,"fullname":"yuze li","user":"llyyzzz","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17566.md","query":{}}">
CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
Published on Aug 18
· Submitted by Fuchen on Aug 19 Abstract
A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence.
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.17566 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.