Skip to content
AI HOT (Curated Pool)

vLLM Releases vllm-metal v0.28.0: Concurrent Serving on Apple Silicon

vLLM 发布 vllm-metal v0.28.0:在 Apple Silicon 上支持并发推理服务

vllm-metal brings vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, tackling high TTFT and memory growth under concurrent local requests. It reuses mlx_lm layers and replaces attention with a custom Metal kernel. v0.28.0 aligns versioning with upstream vLLM, adds batched MTP, GGUF and hybrid model support, and faster M5 prefill. Install via Homebrew; the server speaks the OpenAI API so tools like Claude Code can connect directly. A memory budget set by --gpu-memory-utilization caps the KV cache after a warmup pass, queuing requests that exceed it.

Why it matters: vLLM ports its mature serving stack to Apple Silicon, solving the real pain point of local concurrency with concrete technical details and perf numbers. Not a new model launch, and impact is limited to the Mac ecosystem, so it stays at 78.

Read the original ↗Export Markdown