Skip to content
Hacker News front page

LAION releases BVD: 10M hours of open video data for multimodal pretraining

Laion Big Video Dataset

LAION released BVD, an open dataset with 1.3B video URLs from CommonCrawl, 80M downloaded videos, and 10M total hours. It uses scene detection to create clips with synthetic video and audio captions for multimodal pretraining. ViCLIP models trained on it beat the InternVid baseline by up to 2.1%; CLAP audio models match uncurated audio sets; CLIP trained on 300M extracted frames shows strong image-text retrieval. The release is research-only, non-commercial, and the team flags potential biases and copyright concerns.

Why it matters: LAION drops BVD, a 10M-hour video dataset with 80M videos and a pretrained ViCLIP model that beats InternVid on video-text benchmarks. H and K both hit—scale is clickable, numbers are concrete. Capped at 72 because LAION isn't a model vendor, so R is weak; infra people will ca...

Read the original ↗Export Markdown