vLLM adds distortion-free Gumbel-max text watermarking with weight-free detection
vLLM 新增基于 Gumbel-max 的无失真文本水印功能
vLLM ships text watermarking based on the Gumbel-max trick, embedding a detectable signal during sampling without changing the output distribution. Generation adds only a hash lookup per token; detection needs only the secret key and tokenizer—no model weights or logits. Qwen3.5-27B benchmarks show GSM8K 93.0% vs 94.2% and MBPP 79.2% vs 77.2% with overlapping error bars, so quality holds. The post mentions a dual-key design to resist collusion but doesn't detail key-management practices.
Why it matters: vLLM adds a paper-backed distortion-free watermarking feature — useful for teams doing model serving and compliance. Score capped here because it's infrastructure, not a model capability leap, and the post doesn't disclose latency numbers or detection accuracy.