Skip to content
AI HOT (Curated Pool)

SGLang introduces Weight Cache Daemon for sub-second engine restarts

SGLang 推出 Weight Cache Daemon,实现亚秒级引擎重启

SGLang released Weight Cache Daemon, a persistent GPU process that holds post-quantized model weights in GPU memory and serves them to new engine instances via CUDA IPC zero-copy mapping. On the Ling-2.6-1T FP8 model, weight loading dropped from ~495s to ~0.63s, a ~785× speedup, and total startup fell from 8.8 minutes to 0.528 minutes. The daemon also enables multi-instance weight sharing on the same GPU, active-standby failover in under 1 second, and multi-node support. This is phase one of SGLang's Fast Engine Recovery Framework, targeting sub-10-second cold restarts for production LLM serving.

Why it matters: SGLang's Weight Cache Daemon tackles a real pain point in inference serving: the wait time for weight reloading on engine restart. The CUDA IPC zero-copy approach is clean and backed by Ling-2.6-1T FP8 benchmarks. Score capped because the audience is narrow — directly valuable...

Read the original ↗Export Markdown