Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.</p>\n","updatedAt":"2026-08-21T08:14:33.445Z","author":{"_id":"650817c22a9cebcc9b1cc25d","avatarUrl":"/avatars/40bdf83680b7889ce83a13b01454fcff.svg","fullname":"Yuheng Huang","name":"yuhenghuang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8815544247627258},"editors":["yuhenghuang"],"editorAvatarUrls":["/avatars/40bdf83680b7889ce83a13b01454fcff.svg"],"reactions":[],"isReport":false}},{"id":"6a88092796a608aba0300638","author":{"_id":"650817c22a9cebcc9b1cc25d","avatarUrl":"/avatars/40bdf83680b7889ce83a13b01454fcff.svg","fullname":"Yuheng Huang","name":"yuhenghuang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-21T08:15:35.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This work is supported by Infinimind (https://infinimind.io/en/company/news/2026/narubench-release)","html":"<p>This work is supported by Infinimind (<a href=\"https://infinimind.io/en/company/news/2026/narubench-release\" rel=\"nofollow\">https://infinimind.io/en/company/news/2026/narubench-release</a>)</p>\n","updatedAt":"2026-08-21T08:15:35.665Z","author":{"_id":"650817c22a9cebcc9b1cc25d","avatarUrl":"/avatars/40bdf83680b7889ce83a13b01454fcff.svg","fullname":"Yuheng Huang","name":"yuhenghuang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8653247356414795},"editors":["yuhenghuang"],"editorAvatarUrls":["/avatars/40bdf83680b7889ce83a13b01454fcff.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.13210","authors":[{"_id":"6a8808b689e517cbfd75dd59","user":{"_id":"650817c22a9cebcc9b1cc25d","avatarUrl":"/avatars/40bdf83680b7889ce83a13b01454fcff.svg","isPro":false,"fullname":"Yuheng Huang","user":"yuhenghuang","type":"user","name":"yuhenghuang"},"name":"Yuheng Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-21T08:45:05.302Z","hidden":false},{"_id":"6a8808b689e517cbfd75dd5a","name":"Jianlang Chen","hidden":false},{"_id":"6a8808b689e517cbfd75dd5b","name":"Jiayang Song","hidden":false},{"_id":"6a8808b689e517cbfd75dd5c","name":"Hua Qi","hidden":false},{"_id":"6a8808b689e517cbfd75dd5d","name":"Aza Kai","hidden":false},{"_id":"6a8808b689e517cbfd75dd5e","name":"Vincent Markert","hidden":false},{"_id":"6a8808b689e517cbfd75dd5f","name":"Edison Marrese-Taylor","hidden":false},{"_id":"6a8808b689e517cbfd75dd60","name":"Jianjun Zhao","hidden":false},{"_id":"6a8808b689e517cbfd75dd61","name":"Lei Ma","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video","submittedOnDailyBy":{"_id":"650817c22a9cebcc9b1cc25d","avatarUrl":"/avatars/40bdf83680b7889ce83a13b01454fcff.svg","isPro":false,"fullname":"Yuheng Huang","user":"yuhenghuang","type":"user","name":"yuhenghuang"},"summary":"Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.","upvotes":2,"discussionId":"6a8808b689e517cbfd75dd62","projectPage":"https://ma-labo.github.io/naru/","githubRepo":"https://github.com/infinimind-inc/naru_benchmark","githubRepoAddedBy":"user","ai_summary":"NARU is a Japanese long-form video benchmark evaluating narrative evolution and cultural reasoning through a hierarchical annotation pipeline and extensive native-speaker verification.","ai_keywords":["long-form video understanding","narrative evolution","cultural reasoning","hierarchical memory-based annotation","task-oriented synthesis","shortcut removal","MLLMs"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"6a0fbe68162dc5d32a0057a0","name":"utokyo-ai","fullname":"The University of Tokyo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63a286c7f30c4642278ed11a/jNLYa73JGG4k3ywpoD_9k.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"650817c22a9cebcc9b1cc25d","avatarUrl":"/avatars/40bdf83680b7889ce83a13b01454fcff.svg","isPro":false,"fullname":"Yuheng Huang","user":"yuhenghuang","type":"user"},{"_id":"6426722fc71d90951ea6e4e1","avatarUrl":"/avatars/d7b92c5ce57e9b4fb683745052930b4d.svg","isPro":false,"fullname":"Jiayang Song","user":"sjywdxs","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a0fbe68162dc5d32a0057a0","name":"utokyo-ai","fullname":"The University of Tokyo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63a286c7f30c4642278ed11a/jNLYa73JGG4k3ywpoD_9k.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.13210.md","query":{}}">
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Abstract
NARU is a Japanese long-form video benchmark evaluating narrative evolution and cultural reasoning through a hierarchical annotation pipeline and extensive native-speaker verification.
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.
Community
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.13210 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.13210 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.13210 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.