I would not treat this as a clean benchmark, but the reported shape is exactly what local inference people care about. The summary gives hard hooks: two ASUS GX10 DGX Spark nodes, vLLM, DeepSeek-V4-Flash, TP=2 over RoCE, fp8 KV cache, 256K context, 1,680 prefill tokens/s, 39.8 decode tokens/s with MTP=2, and roughly 1M tokens of safe KV. Reddit 403 blocks the body, so the screenshot, prompt mix, batch size, sampling settings, and reproducibility are not verified.
The spicy part is not the 39.8 tok/s decode. It is the 256K prefill number plus the claimed 1M KV headroom on a tiny two-node setup. That lands right on long-document agents, repo-scale code work, and tool-trace replay. Don’t compare it with H100 fleets; compare it with scrappy hosted inference stacks charging for mediocre long-context latency.