Skip to content
AI HOT (Curated Pool)

Tongyi Lab releases Wan-Streamer v0.2 with 550ms end-to-end latency

通义实验室发布 Wan-Streamer v0.2,端到端响应延迟仅 550ms

Tongyi Lab packed audio, vision, speech, and acting into a single Transformer with 550ms end-to-end latency. v0.2 bumps output resolution from 192×336 to 640×368 at 25 FPS, using a Thinker-Performer dual-path architecture to balance quality and speed. The post doesn't disclose parameter count, training data, or release timeline.

Why it matters: Tongyi Lab's Wan-Streamer v0.2 packs multimodal generation into a single model with hard numbers — 550ms latency and a resolution bump — hitting H and K. R is missing because there's no hook for the broader crowd to latch onto, like a killer use case or a head-to-head win agai...

Read the original ↗Export Markdown