Skip to content
r/LocalLLaMA

meituan-longcat/LongCat-Video-Avatar-1.5 on Hugging Face

meituan-longcat/LongCat-Video-Avatar-1.5 · Hugging Face

Meituan LongCat released LongCat-Video-Avatar-1.5 on Hugging Face, supporting AT2V, ATI2V, and video continuation while replacing Wav2Vec2 with Whisper-Large and using DMD2 distillation to reduce inference to 8 NFE; the model weights are released under the MIT License.

Why it matters: HKR-H/K/R all pass: open MIT video-avatar weights plus 8 NFE inference give local multimodal builders real signal. This is a mid-weight open-source model update, not an 85+ same-day industry event.

Read the original ↗Export Markdown