Recent advances in Omni-LLMs are paving the way for real-time video assistant applications, where models constantly perceive the environment and guide users to achieve certain goals through multi-turn conversations. However, evaluations under these assistant-style interaction scenarios are still challenging. OmniAssistBench aims at addressing this challenge by proposing an annotation pipeline which allows annotators to build test samples from existing Internet videos.</p>\n","updatedAt":"2026-08-24T04:10:02.922Z","author":{"_id":"653b268cd1041ca9188954da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/653b268cd1041ca9188954da/5Qy0fyQdRCIUUAzV6fg1k.png","fullname":"Chaoyou Fu","name":"BradyFU","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9036334753036499},"editors":["BradyFU"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/653b268cd1041ca9188954da/5Qy0fyQdRCIUUAzV6fg1k.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.21360","authors":[{"_id":"6a8b9bdf3d26296ea309189e","name":"Xianyun Sun","hidden":false},{"_id":"6a8b9bdf3d26296ea309189f","name":"Chaoyou Fu","hidden":false},{"_id":"6a8b9bdf3d26296ea30918a0","name":"Zhengye Zhang","hidden":false},{"_id":"6a8b9bdf3d26296ea30918a1","name":"Feiyang Duan","hidden":false},{"_id":"6a8b9bdf3d26296ea30918a2","name":"Qingyuan Cao","hidden":false},{"_id":"6a8b9bdf3d26296ea30918a3","name":"Yonghui Niu","hidden":false},{"_id":"6a8b9bdf3d26296ea30918a4","name":"Sihang Yuan","hidden":false},{"_id":"6a8b9bdf3d26296ea30918a5","name":"Ge Zhang","hidden":false},{"_id":"6a8b9bdf3d26296ea30918a6","name":"Caifeng Shan","hidden":false}],"publishedAt":"2026-08-21T00:00:00.000Z","submittedOnDailyAt":"2026-08-24T00:00:00.000Z","title":"OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs","submittedOnDailyBy":{"_id":"653b268cd1041ca9188954da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/653b268cd1041ca9188954da/5Qy0fyQdRCIUUAzV6fg1k.png","isPro":false,"fullname":"Chaoyou Fu","user":"BradyFU","type":"user","name":"BradyFU"},"summary":"Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to guide users along the exact same routes. Since real interaction videos are rare, we construct the dataset by reverse-engineering existing Internet videos. We deduce logical user goals and segment the videos into multi-turn clips to simulate continuous interactions. This rigorous pipeline required over 1000 expert person-hours to build the dataset. Results show that the proprietary Gemini-3-Pro reaches 66.4 out of the max point of 100, while the open-source Qwen3-Omni-Instruct achieves 51.2. Although current models generally understand user inputs, they frequently provide incorrect or incomplete answers. Specifically, they struggle with visual prompts (e.g., hand gestures), fail to maintain historical context during multi-turn interactions, and fail to delay response until the target event. Results indicate substantial room for improvement before models can become reliable assistants.","upvotes":20,"discussionId":"6a8b9be03d26296ea30918a7","projectPage":"https://xianyunsun.github.io/OmniAssistBench/","githubRepo":"https://github.com/XianyunSun/OmniAssistBench","githubRepoAddedBy":"user","ai_summary":"OmniAssistBench evaluates real-time interactive video assistants by reverse-engineering multi-turn interaction videos, revealing that current omni-modal models struggle with visual prompts, context retention, and timely responses.","ai_keywords":["omni-modal large language models","OmniAssistBench","multi-turn interactions","visual prompts","historical context"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"638f70e8f1256a80d4288555","name":"nanjinguniv","fullname":"Nanjing University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/638f706ef1256a80d42880f9/6M6-JzwJGiLxjIJzvCflf.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69003958f3022902d8220151","avatarUrl":"/avatars/d00400577d78af21f4404afcd3484b85.svg","isPro":false,"fullname":"Xianyun Sun","user":"JontiSun","type":"user"},{"_id":"6a8b9e9f8a2548097a3cd413","avatarUrl":"/avatars/fd078192aa9f79d2c7862bb8351cddd8.svg","isPro":false,"fullname":"WANG WEIXI","user":"kurotsuba","type":"user"},{"_id":"6a8ba14f0e1a59e37d6b8afb","avatarUrl":"/avatars/9583ee6b5f5bce4459bd2fba1ed45a9f.svg","isPro":false,"fullname":"Liu Donglin","user":"KabbageDemon","type":"user"},{"_id":"6628bd00b5d46016492d27e2","avatarUrl":"/avatars/cdbe518c87d7e1bac5dd290019f26bc6.svg","isPro":false,"fullname":"Xinyue Cai","user":"yzlmhzz","type":"user"},{"_id":"64a3df5878edd17f96569747","avatarUrl":"/avatars/e5bdc0d348a9f09b4da05b663e3b710f.svg","isPro":false,"fullname":"lllj","user":"lijiang","type":"user"},{"_id":"653b268cd1041ca9188954da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/653b268cd1041ca9188954da/5Qy0fyQdRCIUUAzV6fg1k.png","isPro":false,"fullname":"Chaoyou Fu","user":"BradyFU","type":"user"},{"_id":"63847adf4d32ba5e1fabef9e","avatarUrl":"/avatars/d6a135ab8606de50b1b351a9922cc16c.svg","isPro":false,"fullname":"Shawn Leo","user":"lxysl","type":"user"},{"_id":"678f682784026ea238b0786d","avatarUrl":"/avatars/f7ca68f366386d13373fac5118a00b5d.svg","isPro":false,"fullname":"牛永辉","user":"summit8848","type":"user"},{"_id":"658cddaf9a1397992a3d7204","avatarUrl":"/avatars/437eb391c146277bb5bc924a5152cdb1.svg","isPro":false,"fullname":"zyzhang","user":"zzygetready","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"67da446beb707e7f71d78d05","avatarUrl":"/avatars/9385bb28f35af7ff6e8473bc20e6d3b9.svg","isPro":false,"fullname":"Flyoung","user":"Flyoung","type":"user"},{"_id":"67e766ddd3f12bae210d76a4","avatarUrl":"/avatars/19cc4d99cf9b603f9e56be142b7e0ddb.svg","isPro":false,"fullname":"Chu Wu","user":"chuwuthw","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"638f70e8f1256a80d4288555","name":"nanjinguniv","fullname":"Nanjing University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/638f706ef1256a80d42880f9/6M6-JzwJGiLxjIJzvCflf.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.21360.md","query":{}}">
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
Abstract
OmniAssistBench evaluates real-time interactive video assistants by reverse-engineering multi-turn interaction videos, revealing that current omni-modal models struggle with visual prompts, context retention, and timely responses.
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to guide users along the exact same routes. Since real interaction videos are rare, we construct the dataset by reverse-engineering existing Internet videos. We deduce logical user goals and segment the videos into multi-turn clips to simulate continuous interactions. This rigorous pipeline required over 1000 expert person-hours to build the dataset. Results show that the proprietary Gemini-3-Pro reaches 66.4 out of the max point of 100, while the open-source Qwen3-Omni-Instruct achieves 51.2. Although current models generally understand user inputs, they frequently provide incorrect or incomplete answers. Specifically, they struggle with visual prompts (e.g., hand gestures), fail to maintain historical context during multi-turn interactions, and fail to delay response until the target event. Results indicate substantial room for improvement before models can become reliable assistants.
Community
Recent advances in Omni-LLMs are paving the way for real-time video assistant applications, where models constantly perceive the environment and guide users to achieve certain goals through multi-turn conversations. However, evaluations under these assistant-style interaction scenarios are still challenging. OmniAssistBench aims at addressing this challenge by proposing an annotation pipeline which allows annotators to build test samples from existing Internet videos.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.21360 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.21360 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.