This is worth a click because it fixes a real pain point: local inference on a Mac works fine for one request, but the moment Claude Code or an agent fires multiple requests, TTFT spikes and memory balloons. vllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, reusing mlx_lm layers underneath and swapping in a custom Metal kernel for attention.
v0.28.0 aligns versioning with upstream vLLM and adds batched MTP, GGUF and hybrid model support, plus faster M5 prefill. Install via Homebrew, and the server speaks the OpenAI API so tools like Claude Code can connect directly.
Two things I'd watch: first, the `--gpu-memory-utilization` memory budget queues requests that exceed it — does that introduce latency jitter under load? Second, since it reuses mlx_lm layers, model support speed depends on mlx_lm's update cadence. Overall, this nudges the Mac from "single-user inference toy" toward "lightweight local server."