7-node ESP32-S3 cluster runs a 0.4B LLM with 1.58-bit ternary quantization
ESP32S3 cluster running 1.58-bit (BitNet) Language model
An open-source project daisy-chains 7 ESP32-S3 boards over SPI to run a 0.4B-parameter language model. The model uses 1.58-bit (BitNet) ternary quantization, so weights are only -1, 0, or 1. The post doesn't disclose inference speed, power draw, or real-world latency, but squeezing a LLM onto cheap microcontrollers is wild.