ChatPaper.aiChatPaper

Audio-Visual Flamingo: 길고 복잡한 비디오를 위한 개방형 오디오-비주얼 인텔리전스

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

July 17, 2026
저자: Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping
cs.AI

초록

본 논문에서는 오디오, 이미지, 장시간 동영상을 공동으로 이해하고 추론할 수 있는 완전히 공개된 최첨단 시청각 대규모 언어 모델(AV-LLM)인 Audio-Visual Flamingo(AV-Flamingo)를 소개합니다. 주로 짧은 클립에 초점을 맞춘 기존 AV-LLM과 달리, AV-Flamingo는 길고 복잡한 실제(시청각) 동영상의 이해와 추론을 위해 설계되었습니다. 이를 위해 세 가지 주요 기여를 합니다: (i) 시간적, 구성적, 교차 모달 시청각 추론을 강조하도록 설계된 약 700만 개의 캡션 및 질문-답변 훈련 인스턴스를 포함하는 대규모 실제 동영상 컬렉션인 Audio-Visual-Skills, (ii) 단거리 인식에서 장기 다중 이벤트 추론으로 모델을 점진적으로 훈련시키는 새로운 3단계 커리큘럼, (iii) 긴 시청각 스트림에서 중간 추론 단계를 타임스탬프에 명시적으로 연결하여 시간적 정렬과 해석 가능성을 향상시키는 추론 프레임워크인 Temporal Audio-Visual Interleaved Chain-of-Thought입니다. 15개 이상의 시청각, 옴니모달, 오디오, 비전 벤치마크에 걸친 광범위한 실험 결과, AV-Flamingo는 유사한 크기의 오픈 모델을 확연한 차이로 능가하며, 특히 길고 복잡한 실제 시청각 이해 및 추론 과제에서 훨씬 더 큰 오픈 가중치 모델 및 폐쇄 모델과도 높은 경쟁력을 유지하고 일부 경우에는 이를 능가함을 보여줍니다. 벤치마크 성능을 넘어, AV-Flamingo는 강력한 실제 유용성을 입증하고 보지 못한 작업으로도 잘 전이되어 견고성과 일반화 능력을 부각시킵니다.
English
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.