Multi-node GPU inference at 30 tok/s over a $20 USB-to-Ethernet adapter
FYI You dont need expensive networking for multi-node gpu. 30t/s laguna Q2_K_XL (39.7GB) on 2x4060+1x4060 using a $20 usb->ethernet.
A Reddit user ran the 39.7 GB laguna Q2_K_XL model across two nodes with three RTX 4060 GPUs, connected only by a direct Ethernet cable. At ubatch 768, generation reached 28.28 tok/s on an 11k-token prompt, with network traffic peaking at 30–70 MB/s. The post includes NCCL+RPC build flags and ubatch comparisons, but does not provide a single-machine dual-GPU baseline for the same model.
Why it matters: A concrete, low-cost multi-node inference experiment: three RTX 4060s over a $20 USB Ethernet adapter hitting 28-30 tok/s, directly challenging the assumption that fast networking is mandatory. Hits all three HKR axes, but it's a community validation rather than a product laun...