WebLLM: Run LLMs directly in your browser, no server needed
WebLLM: high-performance in-browser LLM inference engine
MLC-AI's open-source WebLLM runs LLMs directly in your browser via WebGPU acceleration. It supports Llama, Gemma, and other popular models, achieving near-native speed on consumer GPUs. The catch: first load requires downloading several GB of weights, and memory usage is high. Great for offline assistants and privacy-sensitive use cases, but don't expect it to replace cloud inference.