Skip to content
Hacker News front page

MicroLLM Lab: Run 7 tiny LLMs in your browser with WebGPU

MicroLLM Lab – Try 7 tiny LLM's in the browser

MicroLLM Lab lets you load and run 7 small language models (25M–360M params) directly in your browser via WebGPU, with zero server cost and full data privacy. It's built for edge tasks like query classification, spam filtering, and intent extraction at sub-10ms latency. Models include PetitGPT, SmolLM2, MiniMind2, and GPT-2. You can chat, run objective benchmarks (regex-based), and generate a performance certificate. The post doesn't disclose accuracy on complex tasks—only pass rates on simple pattern checks. Without WebGPU, it falls back to WASM at 8–20 tok/s instead of 100–300 tok/s.

Read the original ↗Export Markdown