Skip to content
r/LocalLLaMA

New MLX LM Server From Apple

A Reddit post says Apple’s MLX LM Server uses continuous batching for concurrent sub-agent requests and supports distributed inference across multiple Macs via Thunderbolt RDMA.

Why it matters: HKR-H/K/R all pass, but the item is based on a Reddit summary and lacks throughput, latency, model-size, or release details. Treat it as a mid-weight Apple/MLX inference update, just above the featured threshold.

Read the original ↗Export Markdown