vLLM Releases vllm-metal v0.28.0: Concurrent Serving on Apple Silicon
What happened
vllm-metal 把 vLLM 的任务调度、分页键值缓存和 OpenAI 兼容接口搬到了苹果芯片上,解决了本地跑模型时多个请求并发导致的响应慢和内存暴涨问题。它复用了 mlx_lm 的模型层,但把注意力计算换成了自定义的 Metal 内核。v0.28.0 版本号跟上游 vLLM 对齐,新增了批量多令牌预测、GGUF 和混合模型支持,并在 M5 芯片...
Coverage
Follow the reports to see the story from different sides.
- AI HOT (Curated Pool)PickvLLM Releases vllm-metal v0.28.0: Concurrent Serving on Apple Silicon
vllm-metal brings vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, tackling high TTFT and memory growth under concurrent local requests. It reuses mlx_lm layers and replaces attention with a custom Metal kernel. v0.28.0 aligns versioning with upstream vLLM, adds batched MTP, GGUF and hybrid model support, and faster M5 prefill. Install via Homebrew; the server speaks the OpenAI API so tools like Claude Code can connect directly. A memory budget set by --gpu-memory-utilization caps the KV cache after a warmup pass, queuing requests that exceed it.