Skip to content
Hacker News front page

Hetzner experiments with an LLM inference API

Hetzner is working on LLM Inference

Hetzner launched an experimental, OpenAI-compatible inference API with a single model: Qwen 3.6 35B FP8. The author measured 153 ms median time-to-first-token and 224 tokens/sec output speed—fast, but the model failed simple arithmetic. There is no billing or SLA yet; Hetzner says it wants to learn about demand, scaling, and load. The more interesting angle is Hetzner's potential to turn spare GPU capacity into a low-margin inference commodity, given its cost-efficient hardware operations. The post does not confirm which GPUs power the service; Hetzner's public GPU lineup includes RTX 4000 SFF (20 GB) and RTX PRO 6000 (96 GB), leaving the hardware question open.

Why it matters: Hetzner dipping into inference is a signal: a European cloud provider moving beyond bare metal into model hosting. The author's benchmarks are solid — latency and throughput look decent — but the model choice (Qwen 3.6 35B FP8) and the arithmetic fail show this is very early-s...

Read the original ↗Export Markdown