This one's worth a look because it's a production inference engine going fully open—MIT or Apache-2.0, kernels included, not the usual "open source but the fast bits are closed" routine.
On a single RTX PRO 6000 running Qwen3.8-27B FP8 with spec decoding off, it beats vLLM across all 13 test cells by 2–19%. Not a blowout, but consistent. Against SGLang it wins 10, loses 2, ties 1. At 32 concurrent clients with 1024/1024 tokens, it hits 1062 tok/s vs vLLM's 958 and SGLang's 844.
I'd throw out the 37x llama.cpp number. A commenter flagged that the llama.cpp benchmark used an unusual KV allocation, and the author didn't push back. The 1.5x figure is the honest baseline.
Limitations are clear: CUDA only, Windows and Linux, validated on Blackwell and Ampere. Ada kernels exist but lack a test board, so the engine refuses to start without an env var. No Mac, ROCm, Vulkan, no tensor parallelism—one model must fit on one GPU. Fine for single-card local setups, irrelevant for multi-GPU users.
The team runs ~300B tokens a year through this thing, so it's not a weekend project. But the benchmark is one card, one model—don't read it as "beats everything everywhere."