视听火烈鸟:面向长复杂视频的开放视听智能
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
July 17, 2026
作者: Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping
cs.AI
摘要
我们提出了音频-视觉火烈鸟(AV-Flamingo),这是一个完全开放的最先进的音频-视觉大型语言模型(AV-LLM),用于对音频、图像和长视频进行联合理解与推理。与先前主要聚焦于短视频片段的AV-LLM不同,AV-Flamingo专为理解和推理长且复杂的现实世界(音频-视觉)视频而设计。为此,我们做出了三项关键贡献:(i)Audio-Visual-Skills,一个大规模的现实世界视频数据集,包含约700万条字幕和问答训练实例,旨在强调时序、组合以及跨模态的音频-视觉推理;(ii)一种新颖的三阶段课程训练方法,逐步将模型从短期感知训练至长期多事件推理;(iii)时序音频-视觉交错思维链,这是一种推理框架,能够将中间推理步骤显式地定位到长音频-视觉流中的时间戳上,从而提升时序对齐能力和可解释性。在超过15个音频-视觉、全模态、音频和视觉基准测试上的大量实验表明,AV-Flamingo以明显优势优于同等规模的开源模型,并且在长且复杂的现实世界音频-视觉理解与推理任务上,与更大规模的开源权重模型和闭源模型相比仍具有很强的竞争力,甚至在某些方面超越它们。除了基准性能之外,AV-Flamingo还展现出强大的现实世界实用性,并能很好地迁移至未见过的任务,突显了其稳健性和泛化能力。
English
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.