Truespar open-sources Paddock, a Rust/C++ inference engine with custom CUDA kernels under MIT/Apache-2.0
We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0)
Truespar open-sourced Paddock, its in-house inference engine written in Rust and C++ with custom CUDA kernels. One binary serves both OpenAI and Anthropic-style APIs, loading GGUF and safetensors. The team runs ~300B tokens/year through it. On an RTX PRO 6000 with Qwen3.8-27B FP8, it beats vLLM by 2–19%, mostly beats SGLang but loses in two cells, and claims 1.5–37x over llama.cpp Q8_0—though a commenter flagged that the llama.cpp benchmark used an unusual KV allocation, so I'd discount the 37x figure. CUDA-only on Windows and Linux, validated on Blackwell and Ampere; Ada kernels exist but lack a validation board, so the engine refuses to start without an env var. No Mac, ROCm, Vulkan, or tensor parallelism—one model per GPU.
Why it matters: Custom CUDA kernels, Rust/C++ engine, internal repo dropped as-is — solid signal for the local LLM community. Backed by 13 benchmark cells, not just speed claims. Held at 72 rather than higher because it's a single Reddit post with no cross-source corroboration yet, and the 30...