Skip to content
Hacker News front page

Jevstiller distills Jev into a local model with a 98% agreement guarantee

Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound

Jevstiller places a small local model in front of the Jev classification API, returning answers in ~15 ms on CPU instead of ~300 ms. It guarantees that, for a chosen target like 98%, the system's overall output agrees with Jev at least that often. It uses a frozen bge-small encoder with a multinomial logistic regression head trained on Jev's full probability distribution, retrained every 2,000 new answers and shadow-tested before promotion. A router combining a confidence threshold and a kNN out-of-distribution scorer decides whether the local model answers or the request is forwarded to Jev. The post derives the agreement formula A = 1 − c·e and explains why a naive confidence threshold breaks the guarantee. Costs, limitations, and a benchmark reproduction command are included.

Why it matters: A model distillation practice with concrete numbers and engineering detail — the 300ms-to-15ms latency drop and the statistical guarantee design are worth reading. But Jev itself has a narrow audience; this reads more like a reference for inference-optimization devs than an in...

Read the original ↗Export Markdown