Video-IFBench evaluates whether video MLLMs can follow complex user instructions with semantic, format, and conditional constraints, revealing substantial gaps beyond standard video understanding accuracy.</p>\n","updatedAt":"2026-08-27T05:16:48.528Z","author":{"_id":"6710be3e6d1b33cf24417e38","avatarUrl":"/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg","fullname":"zpy","name":"zpy777","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7838563323020935},"editors":["zpy777"],"editorAvatarUrls":["/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.25529","authors":[{"_id":"6a8fc8053bd48bb654ea68e8","name":"Hongbo Liu","hidden":false},{"_id":"6a8fc8053bd48bb654ea68e9","name":"Peixian Chen","hidden":false},{"_id":"6a8fc8053bd48bb654ea68ea","name":"Sihan Liu","hidden":false},{"_id":"6a8fc8053bd48bb654ea68eb","name":"Peiyuan Zhang","hidden":false},{"_id":"6a8fc8053bd48bb654ea68ec","name":"Kai Zou","hidden":false},{"_id":"6a8fc8053bd48bb654ea68ed","name":"Dian Zheng","hidden":false},{"_id":"6a8fc8053bd48bb654ea68ee","name":"Xiaoxing Hu","hidden":false},{"_id":"6a8fc8053bd48bb654ea68ef","name":"Yuhao Dong","hidden":false},{"_id":"6a8fc8053bd48bb654ea68f0","name":"Mengdan Zhang","hidden":false},{"_id":"6a8fc8053bd48bb654ea68f1","name":"Yunhang Shen","hidden":false},{"_id":"6a8fc8053bd48bb654ea68f2","name":"Haoyu Cao","hidden":false},{"_id":"6a8fc8053bd48bb654ea68f3","name":"Wei Liu","hidden":false},{"_id":"6a8fc8053bd48bb654ea68f4","name":"Weibo Gu","hidden":false},{"_id":"6a8fc8053bd48bb654ea68f5","name":"Xing Sun","hidden":false},{"_id":"6a8fc8053bd48bb654ea68f6","name":"Shengjie Zhao","hidden":false}],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios","submittedOnDailyBy":{"_id":"6710be3e6d1b33cf24417e38","avatarUrl":"/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg","isPro":false,"fullname":"zpy","user":"zpy777","type":"user","name":"zpy777"},"summary":"Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.","upvotes":13,"discussionId":"6a8fc8053bd48bb654ea68f7","projectPage":"https://alexios-hub.github.io/Video-IFBench/","githubRepo":"https://github.com/Alexios-hub/Video-IFBench","githubRepoAddedBy":"user","ai_summary":"A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints.","ai_keywords":["Multimodal Large Language Models","video understanding","instruction following","Video-IFBench","instruction taxonomy","semantic constraints","format requirements","conditional structures"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6710be3e6d1b33cf24417e38","avatarUrl":"/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg","isPro":false,"fullname":"zpy","user":"zpy777","type":"user"},{"_id":"67e60ae6ac37824273d74389","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/YvPKZ_0gyJnvNwM1zK3JS.png","isPro":true,"fullname":"Dian Zheng","user":"zhengli1013","type":"user"},{"_id":"6809f9a733e15279df7edc5d","avatarUrl":"/avatars/a99889424c9ae78b26484cc368ce178a.svg","isPro":false,"fullname":"ElrosYip","user":"ElrosYip","type":"user"},{"_id":"642a5e5e5673845d9856108f","avatarUrl":"/avatars/72e0303ab6de5330ac518876f4838cb6.svg","isPro":false,"fullname":"Dempsey Wen","user":"DempseyWen","type":"user"},{"_id":"66a85e486419b5b8fcf1a47d","avatarUrl":"/avatars/be20893e45611d5de958540532b79224.svg","isPro":false,"fullname":"liu","user":"ying18","type":"user"},{"_id":"6628766ae7d95899fe6a70e5","avatarUrl":"/avatars/b8f5de38b6350657561147eeab10ec09.svg","isPro":false,"fullname":"huangjiafeng","user":"huangjiafeng","type":"user"},{"_id":"685e3a679b01449f51684163","avatarUrl":"/avatars/decd3c99ec2f9d85b0a91176911fd702.svg","isPro":false,"fullname":"Fxd","user":"fxd0103","type":"user"},{"_id":"691f95388b9cd7dc6b1a52b0","avatarUrl":"/avatars/67717064439877d71482abac5c1df6a9.svg","isPro":false,"fullname":"Aiden Tao","user":"AidenTao","type":"user"},{"_id":"636e19078ba65db4a093a3f4","avatarUrl":"/avatars/287b063b44a022d8576256e80e489c31.svg","isPro":false,"fullname":"alexiosss","user":"Alexislhb","type":"user"},{"_id":"66a27f8cd3449709d69216ce","avatarUrl":"/avatars/71cd4df83a9f086073768c2fc481fc7c.svg","isPro":false,"fullname":"fenfenda","user":"fenfenda","type":"user"},{"_id":"652965773a416e1f2173443b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652965773a416e1f2173443b/y9MB8YgHzbwCXAc4EI9T3.jpeg","isPro":false,"fullname":"Yuhao Dong","user":"THUdyh","type":"user"},{"_id":"6750486e7751c5caa4eda425","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/U5OPyiRtidNwyjcUTOsqg.png","isPro":false,"fullname":"Shilong Dong","user":"FerrisW","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.25529.md","query":{}}">
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Published on Aug 26
· Submitted by zpy on Aug 27 Abstract
A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints.
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.
Community
Video-IFBench evaluates whether video MLLMs can follow complex user instructions with semantic, format, and conditional constraints, revealing substantial gaps beyond standard video understanding accuracy.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.25529 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.25529 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.25529 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.