This caught my eye because it turns decision models into a local runtime you can pull and run like Ollama. You feed it text and a set of typed questions—classification, scoring, yes/no—and it answers in a single forward pass, no token generation. On an RTX 4090, a five-question request to Laya takes 8–10 ms end-to-end, roughly an order of magnitude faster than TypeSafe's hosted Jev API.
Four model families ship out of the box: Laya (322M–421M, English and 100+ languages), decider (Qwen3.5-based, 0.75B–1.9B), nli (zero-shot classifiers, 396M–435M), and gliclass (439M). Weights come straight from Hugging Face, pinned to a commit and sha256-checked. The runtime uses ONNX Runtime and listens on 127.0.0.1 by default.
I'd discount it a bit until we see more. The post doesn't disclose training data or fine-tuning details, and there are no public benchmark numbers to compare against. The decider models run at 155–190 ms on the same 4090—fine for some tasks, but a big gap from Laya's 8 ms. The API compatibility is with TypeSafe's /v1/systemone endpoint, so the ecosystem is narrow for now.
If these models hit your accuracy bar on the classification tasks you actually care about, the real win is keeping sensitive data local and skipping per-call API fees. But you'll need to test precision and generalization yourself.