音視頻火烈鳥:面向長篇幅複雜視頻的開放式音視頻智能
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
July 17, 2026
作者: Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping
cs.AI
摘要
我們提出 Audio-Visual Flamingo(AV-Flamingo),這是一個完全開源、最先進的音訊-視覺大型語言模型(AV-LLM),專為對音訊、影像與長格式影片進行聯合理解與推理而設計。與先前主要專注於短片段音訊-視覺大型語言模型不同,AV-Flamingo 的設計目標是理解與推理長期且複雜的真實世界(音訊-視覺)影片。為此,我們做出三項關鍵貢獻:(i)Audio-Visual-Skills,一個大規模的真實世界影片資料集,包含約 700 萬筆描述與問答訓練實例,特別強調時間性、組合性與跨模態的音訊-視覺推理;(ii)一種新穎的三階段課程學習策略,逐步訓練模型從短時程感知進展到長時程多事件推理;(iii)時序音訊-視覺交錯思維鏈(Temporal Audio-Visual Interleaved Chain-of-Thought),這是一個推理框架,能將中間推理步驟明確對齊到長音訊-視覺串流中的時間戳記,從而改進時間對齊與可解釋性。在超過 15 項音訊-視覺、全模態、音訊與視覺基準測試上的廣泛實驗顯示,AV-Flamingo 明顯優於同等規模的開源模型,並且與更大規模的開放權重模型及封閉模型相比極具競爭力,在某些任務上甚至超越它們,特別是在長期且複雜的真實世界音訊-視覺理解與推理任務上。除了基準測試的表現外,AV-Flamingo 還展現出強大的現實世界實用性與良好的遷移至未見過任務的能力,突顯其穩健性與泛化能力。
English
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.