Skip to content
Hacker News front page

Ternlight: a 7 MB embedding model that runs in-browser with no API calls

Ternlight – 7 MB embedding model that runs in browser (WASM)

Ternlight packs an embedding model into 7 MB, runs on the browser CPU via WASM, and clocks ~5 ms per embed call. It ships as a single npm package with no model download step and no server dependency. A 5 MB mini variant is also available, and the demo shows semantic search over React docs. The post does not disclose model architecture details, training data, or retrieval quality benchmarks beyond the React docs demo. I'd hold off on the 7 MB claim until we see recall on multilingual or long-context tasks.

Why it matters: A 7 MB in-browser embedding model via WASM on CPU with zero server deps has clear practical value for frontend and edge use cases. Score capped at 72 because the post doesn't disclose model architecture, training data, or retrieval quality beyond the React docs demo—those gaps...

Read the original ↗Export Markdown