MicroLLM Lab: Run 7 tiny LLMs in your browser with WebGPU
MicroLLM Lab – Try 7 tiny LLM's in the browser
MicroLLM Lab lets you load and run 7 small language models (25M–360M params) directly in your browser via WebGPU, with zero server cost and full data privacy. It's built for edge tasks like query classification, spam filtering, and intent extraction at sub-10ms latency. Models include PetitGPT, SmolLM2, MiniMind2, and GPT-2. You can chat, run objective benchmarks (regex-based), and generate a performance certificate. The post doesn't disclose accuracy on complex tasks—only pass rates on simple pattern checks. Without WebGPU, it falls back to WASM at 8–20 tok/s instead of 100–300 tok/s.