Skip to content
Trending storyDeveloping

MicroLLM Lab: Run 7 tiny LLMs in your browser with WebGPU

1 report1 sourceupdated 21 hours ago

What happened

Summary

这个项目让你直接在浏览器里加载和运行7个小语言模型(25M到360M参数),靠WebGPU加速,数据不出设备,没有服务器费用。适合做查询分类、垃圾过滤、意图提取这类边缘任务,延迟能压到10毫秒以内。模型包括PetitGPT、SmolLM2、MiniMind2和GPT-2。你可以聊天、跑基准测试(基于正则表达式),还能生成一张性能证书。注意:测试只检查简...

Coverage

Follow the reports to see the story from different sides.

Sep 29
  1. Hacker News front page
    MicroLLM Lab: Run 7 tiny LLMs in your browser with WebGPU

    MicroLLM Lab lets you load and run 7 small language models (25M–360M params) directly in your browser via WebGPU, with zero server cost and full data privacy. It's built for edge tasks like query classification, spam filtering, and intent extraction at sub-10ms latency. Models include PetitGPT, SmolLM2, MiniMind2, and GPT-2. You can chat, run objective benchmarks (regex-based), and generate a performance certificate. The post doesn't disclose accuracy on complex tasks—only pass rates on simple pattern checks. Without WebGPU, it falls back to WASM at 8–20 tok/s instead of 100–300 tok/s.

Heat over time

Not enough continuous observations to draw a trend yet.