Jevstiller distills Jev into a local model with a 98% agreement guarantee
What happened
Jevstiller 在 Jev 分类 API 前面放了一个本地小模型,把每次请求从约 300 毫秒压到 CPU 上约 15 毫秒。它训练的东西很简单:一个冻结的 bge-small 编码器,上面接一个多分类逻辑回归头,用 Jev 返回的完整概率分布来教,而不是只学最终标签。每攒够 2000 条新答案就重训一次,上线前先拿真实流量做影子测试。关键在路由...
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pagePickJevstiller distills Jev into a local model with a 98% agreement guarantee
Jevstiller places a small local model in front of the Jev classification API, returning answers in ~15 ms on CPU instead of ~300 ms. It guarantees that, for a chosen target like 98%, the system's overall output agrees with Jev at least that often. It uses a frozen bge-small encoder with a multinomial logistic regression head trained on Jev's full probability distribution, retrained every 2,000 new answers and shadow-tested before promotion. A router combining a confidence threshold and a kNN out-of-distribution scorer decides whether the local model answers or the request is forwarded to Jev. The post derives the agreement formula A = 1 − c·e and explains why a naive confidence threshold breaks the guarantee. Costs, limitations, and a benchmark reproduction command are included.
Heat over time
Not enough continuous observations to draw a trend yet.