Skip to content
r/LocalLLaMA

Xiaomi serves MiMo V2.5 at 1000–3000 tps with DFlash and Persistent Kernel

Xiaomi is now serving MiMo V2.5 at 1000-3000tps using DFlash & Persistent kernel. DFLash model is out, open-source release promised coming soon

Xiaomi's MiMo V2.5 is live, claiming 1000–3000 tps via DFlash and Persistent Kernel. The DFlash model weights are out, and an open-source release is promised soon. The post body is blocked by Reddit security, so only the headline is available—no details on measured latency, concurrency, or hardware. I'd discount that tps figure: headline peaks usually assume optimal batching, and real single-user throughput is likely lower.

Why it matters: MiMo V2.5's claimed 1000-3000 tps and the two named acceleration mechanisms (DFlash, Persistent Kernel) carry real information density; weights are out and open-source code is promised, directly relevant to local model deployers. Score capped because the Reddit body was blocke...

Read the original ↗Export Markdown