Trending storyDeveloping
7-node ESP32-S3 cluster runs a 0.4B LLM with 1.58-bit ternary quantization
1 report1 sourceupdated 20 hours ago
What happened
Summary
一个开源项目用SPI菊花链把7块ESP32-S3开发板连成集群,跑一个0.4B参数的语言模型。模型用了1.58比特的三值量化(BitNet),权重只存-1、0、1三个值,内存占用极低。但正文没披露推理速度、功耗和实际延迟,所以别急着觉得它能替代GPU。能把大模型塞进几块钱的微控制器里,这个思路本身挺有意思。
Coverage
Follow the reports to see the story from different sides.
Sep 29
- Hacker News front page7-node ESP32-S3 cluster runs a 0.4B LLM with 1.58-bit ternary quantization
An open-source project daisy-chains 7 ESP32-S3 boards over SPI to run a 0.4B-parameter language model. The model uses 1.58-bit (BitNet) ternary quantization, so weights are only -1, 0, or 1. The post doesn't disclose inference speed, power draw, or real-world latency, but squeezing a LLM onto cheap microcontrollers is wild.
Heat over time
Not enough continuous observations to draw a trend yet.