Skip to content
AI HOT · Tips & opinions

vLLM 分离式推理实用指南:何时拆分 prefill 与 decode

vLLM 分离式推理(Disaggregated Serving)实用指南

vLLM 发布分离式推理(Disaggregated Serving)实用指南,说明把 prefill、decode 和 CPU 侧处理拆开能提升满足 TTFT 与 ITL 目标的 goodput,并给出 vLLM v0.30.0 及以上的配置命令。

Read the original ↗Export Markdown