Linum shares its data filtering stack evolution for video model pre-training, from CPU heuristics to RL aesthetic scoring
Getting video models to learn better, faster
Linum is building its open-weight video model v3 and details how its data filtering pipeline evolved since 2024. Early stage used CPU-only traditional CV: PySceneDetect for shot cuts, EAST for OCR, H.264 motion vectors to drop low-motion clips, and Haar cascades to subsample talking heads. By early 2025 they moved to fine-tuned LLMs on GPUs—AutoShot+TransNetV2, PaddleOCR via TensorRT, and Qwen-2-VL-2B for categorical filters. Late 2025 brought RLVR with Qwen-2.5-VL-3B for fine-grained aesthetic scoring (1–4) and WAFT optical flow to catch remaining low-motion long-tail. The post does not disclose v3 release date or model size.
Why it matters: A solid engineering deep-dive on video model data filtering, spanning from 2024 CPU budget hacks to 2025 GPU clusters and RLVR aesthetic filtering. High information density and reusability. Score capped here because the audience is narrow—directly useful for teams building vid...