NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

Yuheng Huang*1, Jianlang Chen*2, Jiayang Song3, Hua Qi1, Aza Kai4, Vincent Markert4, Edison Marrese-Taylor1,4, Jianjun Zhao2, Lei Ma1,5
* Contributed equally to this work
1University of Tokyo, 2Kyushu University, 3Macau University of Science and Technology, 4Infinimind Japan Inc., 5University of Alberta

Introduction

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARUBench, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARUBench consists of 1,481 questions grounded in 155 videos totalling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARUBench offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.

Workflow

Workflow overview of NARUBench.
Workflow overview of NARUBench. Video collection and filtering select 155 long videos from the candidate set. A MLLM-centric pipeline is used to produce taxonomy-aligned evidence for narrative intelligence and cultural understanding annotation. Multiple-choice questions are subsequently generated and refined through a multi-agent pipeline. Verification by 68 native Japanese experts yields the final 1,481 QA items.

Experiment Results

ModelSamplingOverall (%)
N.1N.2N.3N.4Avg.C.1C.2C.3C.4C.5Avg.
Gemini-3-Flash
Google
0.25 FPS83.890.983.978.184.271.369.457.473.269.868.276.2
Gemini-3-Pro
Google
0.25 FPS74.678.676.366.874.160.168.764.269.867.166.070.0
Gemini-2.5-Flash
Google
0.25 FPS51.463.158.150.855.837.849.748.650.348.346.951.4
Qwen3.5-9B
Alibaba
128 Frames34.648.734.934.238.141.349.034.540.941.641.439.8
Qwen3-VL-8B
Alibaba
0.25 FPS35.746.539.835.839.532.240.133.138.932.935.437.4
Qwen2.5-VL-7B
Alibaba
128 Frames25.442.230.626.231.122.440.825.724.228.228.229.7
MiniCPM-o-2.6
OpenBMB
128 Frames23.241.227.431.630.825.937.427.722.128.228.329.6
InternVL3.5
Shanghai AI Lab
64 Frames20.036.425.831.028.332.241.531.835.632.234.631.5

The table reports multiple-choice accuracy across the full benchmark for eight model configurations, broken down into four narrative (N.1–N.4) and five cultural (C.1–C.5) categories with their macro-averages. Performance follows a distinct tiering: proprietary models lead significantly, with Gemini-3-Flash highest at 76.2% overall, while open-source models fall into a lower 29.6–39.8% regime. Category-level results reveal a capability-dependent split: the Gemini models achieve approximately 11 percentage points higher accuracy on narrative tasks than on cultural ones, whereas open-source models show virtually no difference between the two dimensions.

Accuracy vs. number of sampled frames.
Per-dimension gain from denser frame sampling.

The left plot traces overall multiple-choice accuracy on a fixed 500-question subset as the number of sampled frames grows from 8 to 128, and the right plot shows each model’s resulting accuracy gain on the narrative versus cultural dimensions. The three Gemini models improve substantially with larger frame budgets (e.g., Gemini-3-Pro from about 64% to 71%), whereas the open-weight models improve less consistently, most remaining close to 30% even at 128 frames. For every model, narrative gains consistently outpace cultural gains, indicating that narrative errors are largely caused by missing dispersed events that denser sampling directly resolves, while cultural interpretation depends less on visual frequency and more on underlying pragmatic reasoning and domain knowledge.

ModelSamplingNarrative (%)Cultural (%)Overall (%)
N.1N.2N.3N.4C.1C.2C.3C.4C.5
Gemini-3-Flash
Google
0.25 FPS78667275698793858078
Gemini-3-Pro
Google
0.25 FPS77617169658788807274
Gemini-2.5-Flash
Google
0.25 FPS62505865567784687366
Qwen3.5-9B
Alibaba
128 Frames54494750526665605956
Qwen3-VL-8B
Alibaba
0.25 FPS39233436476462395344
Qwen2.5-VL-7B
Alibaba
128 Frames35242728434740284035
MiniCPM-o-2.6
OpenBMB
128 Frames21121516232630153121
InternVL3.5
Shanghai AI Lab
64 Frames40142735434947395138

This table shows the complementary open-ended evaluation on the same 500-question subset: all answer choices are removed, models generate answers directly, and an automated judge scores each response by atomic-fact recall — the fraction of reference facts it covers. Removing the choices preserves the performance hierarchy, with the Gemini family maintaining a substantial advantage (66–78% recall versus 21–56% for open-source models), but the capability profile shifts distinctly. Most strikingly, Sequential/Topical Flow (N.2) — the strongest narrative subcategory under multiple choice — degrades into the weakest narrative category for seven of the eight models, plausibly because multiple-choice options serve as structural scaffolds for temporal organization.

Annotation Pipeline

3. Task-Oriented Annotation
提供された動画チャンクを詳細に分析し、指定されたスキーマに従ってデータを抽出してください。過去の注釈(Recap)は参考情報として提供されます。
日本語で応答してください。

**時間記述の仕様:**
- 出力する全てのタイムスタンプは、現在のチャンク開始時を「0」とした相対時間を使用すること
- 参照用のRecapに含まれる時間は、現在のチャンクに対する相対時間(マイナス値)として記述されていることに留意すること

**出力要件:**
- 指定されたスキーマ構造に従って応答すること
- すべて日本語で応答すること
提供された動画を詳細に分析し、指定されたスキーマに従って**意味的なまとまり(segment)を抽出してください。既存のMLLM注釈(参考情報)を補助として活用してください。

**区間の基準:**
- マクロな構成の抽出: 個々の細かいイベントや会話のやり取りで区切るのではなく、動画全体を章や幕のような大きな構成要素として捉えること
- 包括的なテーマ: 連続する複数のイベントが同一の広範なトピック(例「特定の技術に関する議論」「一つの場所での活動全般」)に属する場合、それらを結合して一つの大きな区間として扱うこと
- 過度な細分化の回避: わずかな場面転換や話者の交代で区切らず、文脈の転換点(topic shift)のみを境界線として設定すること
- タイプの変化による厳密な分割: テーマの継続性にかかわらず、スキーマで定義されたSegmentType(本編からプロモーション、本編から広告など)が切り替わる箇所では、必ず区間を分割すること

**参考情報の扱い:**
- 参考情報として提供される過去の**チャンクの記述**は、3分や5分ごとの機械的・客観的な時間区切りであり、動画の意味的な構造とは無関係です。区間分割の根拠として使用しないでください
- 参考情報内のイベント記述内容を情報源として参照し、実際の区切り位置は動画本編の文脈の変化に基づいて独自に決定してください
- チャンク境界付近のイベントは記述が途切れている(TRUNCATEDと表記)可能性があるため、その不完全性を考慮すること

**出力要件:**
- 指定されたスキーマ構造に従って応答すること
- タイムスタンプは動画内の実際の文脈の切り替わりを正確に反映すること
- すべて日本語で応答すること
提供された動画を詳細に分析し、指定されたスキーマに従って**ナラティブ(narrative)キャプション**を抽出してください。既存のMLLM注釈(参考情報)と文字起こし(利用可能な場合)を補助として活用してください。

**分析指示:**
- 「キャラクターの役(character_roles)」キャラクターを、内面的変容と葛藤を担う主体、一貫した立場から情報伝達や誘導を行う進行役、および場面の文脈補完のみに留まる背景の三種類に抽出すること
- 「叙述展開(narrative_threads)」の抽出と構成
  a. 相互に関連する離散的なイベントの連なりによって形成される、意味的に完結した最小の叙事単位として特定すること
  b. 動画の全体的な流れからこの部分だけを取り出したとしても、背景知識なしに「誰が、いつ、何をして、どうなったか」という一貫したストーリーや意図が成立する独立性を維持すること
  c. 各展開内で、一連のイベントが織りなす文脈、因果関係、および結果を包括的な要約として記述すること

**参考情報の扱い:**
- チャンク境界付近のイベントは記述が途切れている(TRUNCATEDと表記)可能性があるため、その不完全性を考慮すること
- イベントや出来事の文脈を深く理解するために、記述に含まれる対象の詳細情報について、以下の定義リストを補助として活用すること
{Entity Definition Reference}
- 各区間の役割において、以下の定義を補助情報として参照すること
{Segment Type Reference}

**出力要件:**
- 指定されたスキーマ構造に従って応答すること
- すべて日本語で応答すること
提供された動画を詳細に分析し、指定されたスキーマに従って文化的理解データを抽出してください。既存のMLLM注釈と文字起こし(利用可能な場合)を参考コンテキストとして活用してください。

**分析指示:**
1. 動画全体を注意深く視聴し、会話と社会的文脈に特に注意を払う
2. 相槌の使用パターンを識別:「はい」「うん」「ええ」「そうですね」などの機能的役割を分析
3. 空気を読む場面を検出:明示されていない社会的ルールや共有された感情、場の雰囲気を特定
4. 建前と本音を解釈:文字通りの発話と暗示された真意の違いを分析
5. 文化的参照を識別:文化的対象、行動、言及とその意義を説明
6. 感情トーンと対人関係性の推移を評価
  単なる全体の要約ではなく、区間内での雰囲気の変化を捉えること
  a. 安定的: 雰囲気が終始一貫している場合は、リストに単一の評価を含める
  b. 変化あり: 途中で雰囲気が劇的に変化する場合(例:和やかな雑談から緊張した対立へ)は、その推移を表現するために複数の評価を時系列順にリスト化する
  c. 該当なし: イントロ、広告、風景描写のみなど、分析に値する対人関係や感情的文脈が存在しない場合は、無理に出力せず空リストとする

**参考情報の扱い:**
- チャンク境界付近のイベントは記述が途切れている(TRUNCATEDと表記)可能性があるため、その不完全性を考慮すること
- イベントや出来事の文脈を深く理解するために、記述に含まれる対象の詳細情報について、以下の定義リストを補助として活用すること
{Entity Definition Reference}
- 各区間の役割において、以下の定義を補助情報として参照すること
{Segment Type Reference}

**出力要件:**
- 指定されたschema構造に従って応答すること
- タイムスタンプは秒単位で記録すること
- 文化的文脈における適切性と意義を明確に説明すること
- すべて日本語で応答すること
動画の文化的側面、社会的ダイナミクス、感情的ニュアンスを包括的に捉えた文化理解分析を提供してください。

Algorithms

Detailed pseudocode referenced in the paper for video selection and iterative MCQ debiasing.

Incremental semantic diversity filtering used during video selection. Titles and descriptions are embedded with text-embedding-3; a candidate is kept only if its cosine similarity to every selected sample is at most τ = 0.85. The initial seed set size is q = 30.

Algorithm 1 Incremental Semantic Diversity Filtering
  1. Require: Stream of candidates Din, embedding model φ(·), cosine similarity threshold τ, initial set size threshold q
  2. Ensure: Filtered dataset S
  3. S ← LoadCachedState() ▹ Initialize with persisted data
  4. for xDin do
  5. vxφ(x) ▹ Generate normalized embedding
  6. if |S| < q then
  7. SS ∪ {(x, vx)}
  8. else
  9. Let VS be the set of vectors currently in S
  10. σmax ← maxvsVS (vx vs) ▹ Max cosine similarity
  11. if σmaxτ then
  12. SS ∪ {(x, vx)} ▹ Add distinctive sample
  13. UpdateCache(S)
  14. else
  15. continue ▹ Reject redundant sample
  16. end if
  17. end if
  18. end for
  19. return S

Iterative debiasing for generated MCQs. A Blind Solver ensemble answers each item with the video withheld; a Diagnostic Agent attributes any success to a vulnerability and emits a Refine Plan; a Distractor Patching Agent rewrites only the implicated distractor. The stem and correct answer stay fixed. The loop continues until no open item remains vulnerable or the round budget R is exhausted.

Algorithm 2 Iterative Debiasing via the Solver–Critic Loop
  1. Require: MCQ batch Q, Blind-Solver ensemble size N, maximum rounds R
  2. Ensure: Debiased batch Q ▹ stem and correct answer immutable; only distractors change
  3. OQ ▹ items still open for refinement
  4. for r = 1 to R do
  5. for qO do ▹ Blind Solver: video withheld, V = ∅
  6. Cq ← { i : BlindSolveri(q) selects the correct answer } N diversified samples
  7. end for
  8. H ← { qO : Cq = ∅ } ▹ no sample succeeds ⇒ not vulnerable
  9. OO \ H ▹ freeze non-vulnerable items
  10. if O = ∅ then
  11. break ▹ blind accuracy approaches chance
  12. end if
  13. P ← Diagnose({ (q, {reasoning(i) : iCq}) : qO }) ▹ one Refine Plan per vulnerable item
  14. for pP do
  15. e ← Patch(p) ▹ one plan ⇒ one distractor edit
  16. Apply(e, qp) ▹ overwrite the named distractor only
  17. end for
  18. end for
  19. return Q

Each patch names the Refine Plan it addresses and the distractor it targets. Patches that reference an unknown plan or fall outside the plan’s target are rejected and regenerated.