Audio-Visual Flamingo: 長く複雑な動画のためのオープンな音声・映像知能
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
July 17, 2026
著者: Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping
cs.AI
要旨
私たちは、音声、画像、長時間の動画を統合的に理解・推論するための、完全オープンな最先端音声視覚大規模言語モデル(AV-LLM)であるAudio-Visual Flamingo(AV-Flamingo)を提案する。従来のAV-LLMが主に短いクリップに焦点を当てていたのに対し、AV-Flamingoは複雑な実世界の(音声・視覚)動画を長時間にわたって理解・推論できるよう設計されている。これを実現するため、本研究では以下の3つの重要な貢献を行う。(i)約700万件のキャプションと質問応答の訓練インスタンスからなる大規模な実世界動画コレクション「Audio-Visual-Skills」。これは時間的、構成論的、およびクロスモーダルな音声・視覚推論を重視して設計されている。(ii)短期の知覚から長期的なマルチイベント推論へと段階的にモデルを訓練する、新規な3段階カリキュラム。(iii)音声・視覚の長いストリームにおける中間推論ステップを明示的にタイムスタンプに紐づける推論フレームワーク「Temporal Audio-Visual Interleaved Chain-of-Thought」。これにより、時間的整合性と解釈可能性が向上する。15以上の音声・視覚、オムニモーダル、音声、視覚のベンチマークを用いた広範な実験により、AV-Flamingoは同程度のサイズのオープンモデルを明確な差で上回り、特に長く複雑な実世界の音声・視覚理解・推論タスクにおいて、はるかに大規模なオープンウェイトモデルやクローズドモデルと非常に競争力があり、場合によってはそれらを凌駕することを示す。ベンチマーク性能に加えて、AV-Flamingoは実世界での高い有用性を示し、未知のタスクへの転移も良好であり、そのロバスト性と汎化能力を強調している。
English
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.