Hugging Face Daily Papers · · 3 min read

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing</p>\n<p>Project Page: <a href=\"https://coinve200k.github.io/\" rel=\"nofollow\">https://coinve200k.github.io/</a><br>Code: <a href=\"https://github.com/coinve200k/CoinVE-200K\" rel=\"nofollow\">https://github.com/coinve200k/CoinVE-200K</a><br>Dataset: <a href=\"https://huggingface.co/datasets/FireCRT/CoinVE-200K\">https://huggingface.co/datasets/FireCRT/CoinVE-200K</a><br>Model: <a href=\"https://huggingface.co/FireCRT/CoinVE-Edit\">https://huggingface.co/FireCRT/CoinVE-Edit</a><br>Bench: <a href=\"https://huggingface.co/datasets/FireCRT/CoinVE-Bench\">https://huggingface.co/datasets/FireCRT/CoinVE-Bench</a></p>\n","updatedAt":"2026-08-19T06:01:23.225Z","author":{"_id":"6449f2dfeb7db8f70fb990f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png","fullname":"Fuchen","name":"FireCRT","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6024860739707947},"editors":["FireCRT"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17566","authors":[{"_id":"6a851536536bdd3bdd48f7d0","name":"Fuchen Long","hidden":false},{"_id":"6a851536536bdd3bdd48f7d1","name":"Cong Wang","hidden":false},{"_id":"6a851536536bdd3bdd48f7d2","name":"Zitao Gao","hidden":false},{"_id":"6a851536536bdd3bdd48f7d3","name":"Wenhao Zhong","hidden":false},{"_id":"6a851536536bdd3bdd48f7d4","name":"Yu Cheng","hidden":false},{"_id":"6a851536536bdd3bdd48f7d5","name":"Xiaolu Hou","hidden":false},{"_id":"6a851536536bdd3bdd48f7d6","name":"Yan Li","hidden":false},{"_id":"6a851536536bdd3bdd48f7d7","name":"Xiao Cao","hidden":false},{"_id":"6a851536536bdd3bdd48f7d8","name":"Xinlong Sun","hidden":false},{"_id":"6a851536536bdd3bdd48f7d9","name":"Xi Chen","hidden":false},{"_id":"6a851536536bdd3bdd48f7da","name":"Yu Liu","hidden":false}],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing","submittedOnDailyBy":{"_id":"6449f2dfeb7db8f70fb990f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png","isPro":false,"fullname":"Fuchen","user":"FireCRT","type":"user","name":"FireCRT"},"summary":"The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.","upvotes":10,"discussionId":"6a851536536bdd3bdd48f7db","projectPage":"https://coinve200k.github.io/","githubRepo":"https://github.com/coinve200k/CoinVE-200K","githubRepoAddedBy":"user","ai_summary":"A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence.","ai_keywords":["compositional instruction-guided video editing","CoinVE-200K","CoinVE-Bench","CoinVE-Edit","region-aware attention","Wan2.1-T2V-14B","Qwen3-VL-8B-Instruct","temporal consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6449f2dfeb7db8f70fb990f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6449f2dfeb7db8f70fb990f8/limS6x8txJJHCIBDbih-N.png","isPro":false,"fullname":"Fuchen","user":"FireCRT","type":"user"},{"_id":"6556ea0e0f4493529783e7a4","avatarUrl":"/avatars/df404cae1511414ce648469e9f0f0714.svg","isPro":false,"fullname":"PyBigStar","user":"PyBigStar","type":"user"},{"_id":"644b96709ca7117968e75997","avatarUrl":"/avatars/3ba8fbde30feee7f117ec015ce13dc15.svg","isPro":false,"fullname":"BiBiKo_219","user":"BiBiKo","type":"user"},{"_id":"646339704ad7e61e51db904c","avatarUrl":"/avatars/a073ac880bcc3ec5f3a832da4a94ed8e.svg","isPro":false,"fullname":"SkyScraper","user":"Cong723","type":"user"},{"_id":"68834a9283873f346bfc88bc","avatarUrl":"/avatars/4643d7931857c9c3bf9b192ca8d15a65.svg","isPro":false,"fullname":"binyuanhuang","user":"foralex","type":"user"},{"_id":"6878987600ce33ca49b91d3f","avatarUrl":"/avatars/e30e08c2f8bfdb4f00103327dd43c21a.svg","isPro":false,"fullname":"cy","user":"wuucy","type":"user"},{"_id":"6496f5754a3c31df8e3139f6","avatarUrl":"/avatars/cf789d1986f976373c82b2976df4542a.svg","isPro":false,"fullname":"Zhongwei Zhang","user":"zzwustc","type":"user"},{"_id":"686349186139862d1f7d1be3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/YTthC2OhhUAu5dQ46EjlH.png","isPro":false,"fullname":"wehaozhong","user":"wenhaozhong","type":"user"},{"_id":"65f1642abf8631e7dae1db4e","avatarUrl":"/avatars/4d6256b94abeb2ff6894b1f9a2c11db7.svg","isPro":false,"fullname":"yucheng","user":"Yumic","type":"user"},{"_id":"6678dc9fa845e4470fab99fc","avatarUrl":"/avatars/c10f1ec6be0eedfb9cf07a6d2cb700a3.svg","isPro":false,"fullname":"yuze li","user":"llyyzzz","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17566.md","query":{}}">
Papers
arxiv:2608.17566

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

Published on Aug 18
· Submitted by
Fuchen
on Aug 19
Authors:
,

Abstract

A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence.

The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.

Community

Paper submitter about 2 hours ago

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

Project Page: https://coinve200k.github.io/
Code: https://github.com/coinve200k/CoinVE-200K
Dataset: https://huggingface.co/datasets/FireCRT/CoinVE-200K
Model: https://huggingface.co/FireCRT/CoinVE-Edit
Bench: https://huggingface.co/datasets/FireCRT/CoinVE-Bench

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.17566
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.17566 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers