Skip to content
r/LocalLLaMA

Xiaomi claims 1,000+ tps on a 1T model using a standard 8-GPU server

Xiaomi just claimed 1,000+ tps on a 1T model using a standard 8-GPU server

Xiaomi MiMo claims MiMo-V2.5-Pro UltraSpeed runs a 1T-parameter MoE model above 1,000 output tokens per second on one standard 8-GPU node; the post does not disclose the GPU model, batch settings, or reproducible configuration.

Why it matters: HKR-H/K/R all pass: the 1T MoE and 1,000+ tps claim is a strong inference-cost hook. Kept below P1 because the post lacks GPU model, batch size, quantization, and reproducible setup.

Read the original ↗Export Markdown