Mesh LLM pools GPUs across machines into one OpenAI-compatible inference mesh
Mesh LLM: distributed AI computing on iroh
The n0 team open-sourced Mesh LLM, which pools GPUs from multiple ordinary machines to run large models and exposes an OpenAI-compatible API at localhost:9337/v1. Requests can run locally, route to a peer with the model loaded, or split a model across machines in a layer-wise pipeline. Underneath, iroh handles NAT traversal and direct QUIC connections with no central server. Over 40 models are included, from 0.5B to 235B MoE. The post doesn't disclose inference latency or throughput numbers, so I'd hold off on production assumptions.
Why it matters: n0 team open-sourced a way to pool GPUs across ordinary machines for LLM inference over iroh's P2P QUIC transport, no central server. The mechanism is well-explained and appeals to cost-conscious, privacy-sensitive teams. Score held at 72 because it's a new project with no rea...