Skip to content
Trending storyPast story

vLLM Releases vllm-metal v0.28.0: Concurrent Serving on Apple Silicon

1 report1 sourceupdated 6 days ago

What happened

Summary

vllm-metal 把 vLLM 的任务调度、分页键值缓存和 OpenAI 兼容接口搬到了苹果芯片上,解决了本地跑模型时多个请求并发导致的响应慢和内存暴涨问题。它复用了 mlx_lm 的模型层,但把注意力计算换成了自定义的 Metal 内核。v0.28.0 版本号跟上游 vLLM 对齐,新增了批量多令牌预测、GGUF 和混合模型支持,并在 M5 芯片...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. AI HOT (Curated Pool)Pick
    vLLM Releases vllm-metal v0.28.0: Concurrent Serving on Apple Silicon

    vllm-metal brings vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, tackling high TTFT and memory growth under concurrent local requests. It reuses mlx_lm layers and replaces attention with a custom Metal kernel. v0.28.0 aligns versioning with upstream vLLM, adds batched MTP, GGUF and hybrid model support, and faster M5 prefill. Install via Homebrew; the server speaks the OpenAI API so tools like Claude Code can connect directly. A memory budget set by --gpu-memory-utilization caps the KV cache after a warmup pass, queuing requests that exceed it.