‹ 返回 2026-07-21

视听火烈鸟:为长且复杂的视频提供开放的视听智能

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

▲ 5 💬 1 2026-07-21

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping

摘要

我们提出了Audio-Visual Flamingo(AV-Flamingo),这是一种完全开放式的先进音频-视觉大语言模型(AV-LLM),能够用于对音频、图像和长视频进行联合理解与推理。与那些主要处理短片段的先前AV-LLM不同,AV-Flamingo旨在处理长且复杂的现实场景中的音频-视觉视频。为了实现这一目标,我们做出了三项关键贡献:(i) Audio-Visual-Skills:这是一个包含约700万条带有说明和问答样本的大规模现实视频数据集,旨在促进时间、构图以及跨模态的音频-视觉推理能力;(ii)一种创新的三阶段训练体系,该体系从短期感知逐步过渡到长期多事件推理;以及(iii)时间性音频-视觉混合思维链推理框架,该框架将中间推理步骤与长音频-视觉流中的时间戳相关联,从而提升时间上的一致性与可解释性。在15多个音频-视觉、全模态、音频和视觉测试基准上的大量实验表明,AV-Flamingo在性能上明显优于同规模的开源模型,并且与一些更大规模的开源或封闭模型相比也具有竞争力,尤其是在处理长且复杂的现实场景中的音频-视觉理解与推理任务时。除了测试基准性能之外,AV-Flamingo还具备强大的现实应用价值,能够很好地应用于未知任务,这凸显了它的鲁棒性和泛化能力。

English Abstract

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.