Skip to content
AI HOT (Curated Pool)

NVIDIA releases Audio-Visual Flamingo: an open AV-LLM for long, complex videos

NVIDIA 发布 Audio-Visual Flamingo:面向长视频的开放音频-视觉大语言模型

NVIDIA fully open-sourced Audio-Visual Flamingo, a model that jointly understands and reasons over audio, images, and long-form videos. Unlike most AV-LLMs that focus on short clips, it targets real-world long videos. It was trained on ~7M timestamped caption and QA instances using a three-stage curriculum that moves from short-range perception to long-horizon multi-event reasoning. A reasoning framework called TAVI-CoT grounds intermediate steps to specific timestamps. Across 15+ audio, vision, and multimodal benchmarks, AV-Flamingo clearly beats similarly sized open models and matches or surpasses larger closed models on long, complex audio-visual tasks.

Why it matters: NVIDIA open-sourced a model that handles both visual and audio streams in long videos, with 7M timestamped samples and a three-stage training recipe — real new info for video-understanding builders. Score stays at the featured threshold because it's a pure research release wit...

Read the original ↗Export Markdown