TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Google DeepMind released TIPSv2 with 3 pretraining changes for CVPR 2026. iBOT++ applies self-distillation to masked and visible patches, adding 14.1 mIoU on ADE150; Head-only EMA cuts training parameters by 42%. The key signal is visible-token supervision, not a larger teacher model.
Why it matters: HKR-K is strong: ADE150 gains 14.1 mIoU and trainable params drop 42%. HKR-H/R pass, but this is still a VLM research release, not a same-day model launch.