NVIDIA releases Audio-Visual Flamingo: an open AV-LLM for long, complex videos
NVIDIA 发布 Audio-Visual Flamingo:面向长视频的开放音频-视觉大语言模型
NVIDIA fully open-sourced Audio-Visual Flamingo, a model that jointly understands and reasons over audio, images, and long-form videos. Unlike most AV-LLMs that focus on short clips, it targets real-world long videos. It was trained on ~7M timestamped caption and QA instances using a three-stage curriculum that moves from short-range perception to long-horizon multi-event reasoning. A reasoning framework called TAVI-CoT grounds intermediate steps to specific timestamps. Across 15+ audio, vision, and multimodal benchmarks, AV-Flamingo clearly beats similarly sized open models and matches or surpasses larger closed models on long, complex audio-visual tasks.
Why it matters: NVIDIA open-sourced a model that handles both visual and audio streams in long videos, with 7M timestamped samples and a three-stage training recipe — real new info for video-understanding builders. Score stays at the featured threshold because it's a pure research release wit...