| Model | Sampling | Overall (%) | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| N.1 | N.2 | N.3 | N.4 | Avg. | C.1 | C.2 | C.3 | C.4 | C.5 | Avg. | |||
| Gemini-3-Flash Google | 0.25 FPS | 83.8 | 90.9 | 83.9 | 78.1 | 84.2 | 71.3 | 69.4 | 57.4 | 73.2 | 69.8 | 68.2 | 76.2 |
| Gemini-3-Pro Google | 0.25 FPS | 74.6 | 78.6 | 76.3 | 66.8 | 74.1 | 60.1 | 68.7 | 64.2 | 69.8 | 67.1 | 66.0 | 70.0 |
| Gemini-2.5-Flash Google | 0.25 FPS | 51.4 | 63.1 | 58.1 | 50.8 | 55.8 | 37.8 | 49.7 | 48.6 | 50.3 | 48.3 | 46.9 | 51.4 |
| Qwen3.5-9B Alibaba | 128 Frames | 34.6 | 48.7 | 34.9 | 34.2 | 38.1 | 41.3 | 49.0 | 34.5 | 40.9 | 41.6 | 41.4 | 39.8 |
| Qwen3-VL-8B Alibaba | 0.25 FPS | 35.7 | 46.5 | 39.8 | 35.8 | 39.5 | 32.2 | 40.1 | 33.1 | 38.9 | 32.9 | 35.4 | 37.4 |
| Qwen2.5-VL-7B Alibaba | 128 Frames | 25.4 | 42.2 | 30.6 | 26.2 | 31.1 | 22.4 | 40.8 | 25.7 | 24.2 | 28.2 | 28.2 | 29.7 |
| MiniCPM-o-2.6 OpenBMB | 128 Frames | 23.2 | 41.2 | 27.4 | 31.6 | 30.8 | 25.9 | 37.4 | 27.7 | 22.1 | 28.2 | 28.3 | 29.6 |
| InternVL3.5 Shanghai AI Lab | 64 Frames | 20.0 | 36.4 | 25.8 | 31.0 | 28.3 | 32.2 | 41.5 | 31.8 | 35.6 | 32.2 | 34.6 | 31.5 |
The table reports multiple-choice accuracy across the full benchmark for eight model configurations, broken down into four narrative (N.1–N.4) and five cultural (C.1–C.5) categories with their macro-averages. Performance follows a distinct tiering: proprietary models lead significantly, with Gemini-3-Flash highest at 76.2% overall, while open-source models fall into a lower 29.6–39.8% regime. Category-level results reveal a capability-dependent split: the Gemini models achieve approximately 11 percentage points higher accuracy on narrative tasks than on cultural ones, whereas open-source models show virtually no difference between the two dimensions.
The left plot traces overall multiple-choice accuracy on a fixed 500-question subset as the number of sampled frames grows from 8 to 128, and the right plot shows each model’s resulting accuracy gain on the narrative versus cultural dimensions. The three Gemini models improve substantially with larger frame budgets (e.g., Gemini-3-Pro from about 64% to 71%), whereas the open-weight models improve less consistently, most remaining close to 30% even at 128 frames. For every model, narrative gains consistently outpace cultural gains, indicating that narrative errors are largely caused by missing dispersed events that denser sampling directly resolves, while cultural interpretation depends less on visual frequency and more on underlying pragmatic reasoning and domain knowledge.
| Model | Sampling | Narrative (%) | Cultural (%) | Overall (%) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| N.1 | N.2 | N.3 | N.4 | C.1 | C.2 | C.3 | C.4 | C.5 | |||
| Gemini-3-Flash Google | 0.25 FPS | 78 | 66 | 72 | 75 | 69 | 87 | 93 | 85 | 80 | 78 |
| Gemini-3-Pro Google | 0.25 FPS | 77 | 61 | 71 | 69 | 65 | 87 | 88 | 80 | 72 | 74 |
| Gemini-2.5-Flash Google | 0.25 FPS | 62 | 50 | 58 | 65 | 56 | 77 | 84 | 68 | 73 | 66 |
| Qwen3.5-9B Alibaba | 128 Frames | 54 | 49 | 47 | 50 | 52 | 66 | 65 | 60 | 59 | 56 |
| Qwen3-VL-8B Alibaba | 0.25 FPS | 39 | 23 | 34 | 36 | 47 | 64 | 62 | 39 | 53 | 44 |
| Qwen2.5-VL-7B Alibaba | 128 Frames | 35 | 24 | 27 | 28 | 43 | 47 | 40 | 28 | 40 | 35 |
| MiniCPM-o-2.6 OpenBMB | 128 Frames | 21 | 12 | 15 | 16 | 23 | 26 | 30 | 15 | 31 | 21 |
| InternVL3.5 Shanghai AI Lab | 64 Frames | 40 | 14 | 27 | 35 | 43 | 49 | 47 | 39 | 51 | 38 |
This table shows the complementary open-ended evaluation on the same 500-question subset: all answer choices are removed, models generate answers directly, and an automated judge scores each response by atomic-fact recall — the fraction of reference facts it covers. Removing the choices preserves the performance hierarchy, with the Gemini family maintaining a substantial advantage (66–78% recall versus 21–56% for open-source models), but the capability profile shifts distinctly. Most strikingly, Sequential/Topical Flow (N.2) — the strongest narrative subcategory under multiple choice — degrades into the weakest narrative category for seven of the eight models, plausibly because multiple-choice options serve as structural scaffolds for temporal organization.
提供された動画チャンクを詳細に分析し、指定されたスキーマに従ってデータを抽出してください。過去の注釈(Recap)は参考情報として提供されます。
日本語で応答してください。
**時間記述の仕様:**
- 出力する全てのタイムスタンプは、現在のチャンク開始時を「0」とした相対時間を使用すること
- 参照用のRecapに含まれる時間は、現在のチャンクに対する相対時間(マイナス値)として記述されていることに留意すること
**出力要件:**
- 指定されたスキーマ構造に従って応答すること
- すべて日本語で応答すること
提供された動画を詳細に分析し、指定されたスキーマに従って**意味的なまとまり(segment)を抽出してください。既存のMLLM注釈(参考情報)を補助として活用してください。
**区間の基準:**
- マクロな構成の抽出: 個々の細かいイベントや会話のやり取りで区切るのではなく、動画全体を章や幕のような大きな構成要素として捉えること
- 包括的なテーマ: 連続する複数のイベントが同一の広範なトピック(例「特定の技術に関する議論」「一つの場所での活動全般」)に属する場合、それらを結合して一つの大きな区間として扱うこと
- 過度な細分化の回避: わずかな場面転換や話者の交代で区切らず、文脈の転換点(topic shift)のみを境界線として設定すること
- タイプの変化による厳密な分割: テーマの継続性にかかわらず、スキーマで定義されたSegmentType(本編からプロモーション、本編から広告など)が切り替わる箇所では、必ず区間を分割すること
**参考情報の扱い:**
- 参考情報として提供される過去の**チャンクの記述**は、3分や5分ごとの機械的・客観的な時間区切りであり、動画の意味的な構造とは無関係です。区間分割の根拠として使用しないでください
- 参考情報内のイベント記述内容を情報源として参照し、実際の区切り位置は動画本編の文脈の変化に基づいて独自に決定してください
- チャンク境界付近のイベントは記述が途切れている(TRUNCATEDと表記)可能性があるため、その不完全性を考慮すること
**出力要件:**
- 指定されたスキーマ構造に従って応答すること
- タイムスタンプは動画内の実際の文脈の切り替わりを正確に反映すること
- すべて日本語で応答すること
提供された動画を詳細に分析し、指定されたスキーマに従って**ナラティブ(narrative)キャプション**を抽出してください。既存のMLLM注釈(参考情報)と文字起こし(利用可能な場合)を補助として活用してください。
**分析指示:**
- 「キャラクターの役(character_roles)」キャラクターを、内面的変容と葛藤を担う主体、一貫した立場から情報伝達や誘導を行う進行役、および場面の文脈補完のみに留まる背景の三種類に抽出すること
- 「叙述展開(narrative_threads)」の抽出と構成
a. 相互に関連する離散的なイベントの連なりによって形成される、意味的に完結した最小の叙事単位として特定すること
b. 動画の全体的な流れからこの部分だけを取り出したとしても、背景知識なしに「誰が、いつ、何をして、どうなったか」という一貫したストーリーや意図が成立する独立性を維持すること
c. 各展開内で、一連のイベントが織りなす文脈、因果関係、および結果を包括的な要約として記述すること
**参考情報の扱い:**
- チャンク境界付近のイベントは記述が途切れている(TRUNCATEDと表記)可能性があるため、その不完全性を考慮すること
- イベントや出来事の文脈を深く理解するために、記述に含まれる対象の詳細情報について、以下の定義リストを補助として活用すること
{Entity Definition Reference}
- 各区間の役割において、以下の定義を補助情報として参照すること
{Segment Type Reference}
**出力要件:**
- 指定されたスキーマ構造に従って応答すること
- すべて日本語で応答すること
提供された動画を詳細に分析し、指定されたスキーマに従って文化的理解データを抽出してください。既存のMLLM注釈と文字起こし(利用可能な場合)を参考コンテキストとして活用してください。
**分析指示:**
1. 動画全体を注意深く視聴し、会話と社会的文脈に特に注意を払う
2. 相槌の使用パターンを識別:「はい」「うん」「ええ」「そうですね」などの機能的役割を分析
3. 空気を読む場面を検出:明示されていない社会的ルールや共有された感情、場の雰囲気を特定
4. 建前と本音を解釈:文字通りの発話と暗示された真意の違いを分析
5. 文化的参照を識別:文化的対象、行動、言及とその意義を説明
6. 感情トーンと対人関係性の推移を評価
単なる全体の要約ではなく、区間内での雰囲気の変化を捉えること
a. 安定的: 雰囲気が終始一貫している場合は、リストに単一の評価を含める
b. 変化あり: 途中で雰囲気が劇的に変化する場合(例:和やかな雑談から緊張した対立へ)は、その推移を表現するために複数の評価を時系列順にリスト化する
c. 該当なし: イントロ、広告、風景描写のみなど、分析に値する対人関係や感情的文脈が存在しない場合は、無理に出力せず空リストとする
**参考情報の扱い:**
- チャンク境界付近のイベントは記述が途切れている(TRUNCATEDと表記)可能性があるため、その不完全性を考慮すること
- イベントや出来事の文脈を深く理解するために、記述に含まれる対象の詳細情報について、以下の定義リストを補助として活用すること
{Entity Definition Reference}
- 各区間の役割において、以下の定義を補助情報として参照すること
{Segment Type Reference}
**出力要件:**
- 指定されたschema構造に従って応答すること
- タイムスタンプは秒単位で記録すること
- 文化的文脈における適切性と意義を明確に説明すること
- すべて日本語で応答すること
動画の文化的側面、社会的ダイナミクス、感情的ニュアンスを包括的に捉えた文化理解分析を提供してください。
Detailed pseudocode referenced in the paper for video selection and iterative MCQ debiasing.
Incremental semantic diversity filtering used during video selection. Titles and descriptions are embedded with text-embedding-3; a candidate is kept only if its cosine similarity to every selected sample is at most τ = 0.85. The initial seed set size is q = 30.
Iterative debiasing for generated MCQs. A Blind Solver ensemble answers each item with the video withheld; a Diagnostic Agent attributes any success to a vulnerability and emits a Refine Plan; a Distractor Patching Agent rewrites only the implicated distractor. The stem and correct answer stay fixed. The loop continues until no open item remains vulnerable or the round budget R is exhausted.
Each patch names the Refine Plan it addresses and the distractor it targets. Patches that reference an unknown plan or fall outside the plan’s target are rejected and regenerated.