MicroLLM Lab: Run 7 tiny LLMs in your browser with WebGPU
What happened
这个项目让你直接在浏览器里加载和运行7个小语言模型(25M到360M参数),靠WebGPU加速,数据不出设备,没有服务器费用。适合做查询分类、垃圾过滤、意图提取这类边缘任务,延迟能压到10毫秒以内。模型包括PetitGPT、SmolLM2、MiniMind2和GPT-2。你可以聊天、跑基准测试(基于正则表达式),还能生成一张性能证书。注意:测试只检查简...
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageMicroLLM Lab: Run 7 tiny LLMs in your browser with WebGPU
MicroLLM Lab lets you load and run 7 small language models (25M–360M params) directly in your browser via WebGPU, with zero server cost and full data privacy. It's built for edge tasks like query classification, spam filtering, and intent extraction at sub-10ms latency. Models include PetitGPT, SmolLM2, MiniMind2, and GPT-2. You can chat, run objective benchmarks (regex-based), and generate a performance certificate. The post doesn't disclose accuracy on complex tasks—only pass rates on simple pattern checks. Without WebGPU, it falls back to WASM at 8–20 tok/s instead of 100–300 tok/s.
Heat over time
Not enough continuous observations to draw a trend yet.