Hugging Face Daily Papers · · 3 min read

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Code: <a href=\"https://github.com/zfkarl/VideoGAIA\" rel=\"nofollow\">https://github.com/zfkarl/VideoGAIA</a><br>Data: <a href=\"https://huggingface.co/datasets/Karl28/VideoGAIA\">https://huggingface.co/datasets/Karl28/VideoGAIA</a></p>\n","updatedAt":"2026-08-18T03:21:29.302Z","author":{"_id":"6639ad487c0ab4fd9df1dde5","avatarUrl":"/avatars/8cc99f6ed8f8c1b2a14dde797a991a8c.svg","fullname":"Fan Zhang","name":"Karl28","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5769637227058411},"editors":["Karl28"],"editorAvatarUrls":["/avatars/8cc99f6ed8f8c1b2a14dde797a991a8c.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.14718","authors":[{"_id":"6a83cf6e675db694db8cd55a","name":"Fan Zhang","hidden":false},{"_id":"6a83cf6e675db694db8cd55b","name":"Guangming Yao","hidden":false},{"_id":"6a83cf6e675db694db8cd55c","name":"Jinyang Wu","hidden":false},{"_id":"6a83cf6e675db694db8cd55d","name":"Hao Wu","hidden":false},{"_id":"6a83cf6e675db694db8cd55e","name":"Zheng Lian","hidden":false},{"_id":"6a83cf6e675db694db8cd55f","name":"Xinyu Geng","hidden":false},{"_id":"6a83cf6e675db694db8cd560","name":"Jingdong Chen","hidden":false},{"_id":"6a83cf6e675db694db8cd561","name":"Yi Yuan","hidden":false},{"_id":"6a83cf6e675db694db8cd562","name":"Pheng-Ann Heng","hidden":false}],"publishedAt":"2026-08-12T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding","submittedOnDailyBy":{"_id":"6639ad487c0ab4fd9df1dde5","avatarUrl":"/avatars/8cc99f6ed8f8c1b2a14dde797a991a8c.svg","isPro":false,"fullname":"Fan Zhang","user":"Karl28","type":"user","name":"Karl28"},"summary":"Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.","upvotes":8,"discussionId":"6a83cf6e675db694db8cd563","githubRepo":"https://github.com/zfkarl/VideoGAIA","githubRepoAddedBy":"user","ai_summary":"VideoGAIA introduces a multi-turn, tool-augmented benchmark that evaluates agentic video understanding for advanced multimodal models through complex real-world tasks.","ai_keywords":["multimodal large language models","agentic video understanding","tool-augmented interaction","VideoGAIA","multi-turn interaction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6390c6fdd00f25601f445cd4","name":"CUHK-CSE","fullname":"The Chinese University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/621f2eb36e152b56a7cf0248/o8RRAczRjfNEzq70GzUwQ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6639ad487c0ab4fd9df1dde5","avatarUrl":"/avatars/8cc99f6ed8f8c1b2a14dde797a991a8c.svg","isPro":false,"fullname":"Fan Zhang","user":"Karl28","type":"user"},{"_id":"6635c12178f0395f0b7f3602","avatarUrl":"/avatars/c6a86828266d7d3259b0c19b916227f7.svg","isPro":false,"fullname":"wu","user":"23davin","type":"user"},{"_id":"640bfda15d36b591f04465f6","avatarUrl":"/avatars/aa1467172cdd34c56bd5a9d1933b114a.svg","isPro":false,"fullname":"Gu","user":"yuanqi","type":"user"},{"_id":"695fbbf71aaeaadac0d1b294","avatarUrl":"/avatars/96833150641b1eb90d2a325bdf6b8347.svg","isPro":false,"fullname":"Zhou","user":"LilyZhou1","type":"user"},{"_id":"682b22ebac526172e1b4ed1b","avatarUrl":"/avatars/a9e486bf72d27013e6c1903b64a7754c.svg","isPro":false,"fullname":"Geng Xinyu","user":"Ornamentt","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"61d78857a21a9b49c7e8e4a9","avatarUrl":"/avatars/c7e7f84cad775be2d13fab8530bf21f5.svg","isPro":false,"fullname":"Yifan Du","user":"Richard1999","type":"user"},{"_id":"6606d72ceea08fc29683bfd5","avatarUrl":"/avatars/76bd32d49b05330c6b328f4a7ad9baf0.svg","isPro":false,"fullname":"shuo yang","user":"shuo-yan","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6390c6fdd00f25601f445cd4","name":"CUHK-CSE","fullname":"The Chinese University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/621f2eb36e152b56a7cf0248/o8RRAczRjfNEzq70GzUwQ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.14718.md","query":{}}">
Papers
arxiv:2608.14718

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

Published on Aug 12
· Submitted by
Fan Zhang
on Aug 18
Authors:
,

Abstract

VideoGAIA introduces a multi-turn, tool-augmented benchmark that evaluates agentic video understanding for advanced multimodal models through complex real-world tasks.

Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasingly saturated and insufficient for assessing the intelligence of advanced MLLMs. Towards this end, we introduce VideoGAIA, an agentic video understanding benchmark for general artificial intelligence (AI) assistants. Moving beyond one-shot video question answering, VideoGAIA formulates video understanding as a multi-turn, tool-augmented interaction process, where models must iteratively perceive videos, invoke external tools, gather complementary information, and integrate multimodal evidence across turns. VideoGAIA contains 271 model-human co-designed tasks covering diverse and complex real-world scenarios. Each video-question-answer instance is independently verified by three human experts to ensure both correctness and appropriate difficulty. All evaluated MLLMs, including frontier models such as GPT-5.5 and Kimi-K3, achieve less than 60% accuracy on VideoGAIA, highlighting its value as a high-quality and timely benchmark for evaluating next-generation MLLMs. We hope that VideoGAIA will facilitate the transition from conventional video understanding toward agentic video understanding.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.14718
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.14718 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.14718 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.14718 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers