Ollaya runs open-source decision models locally with single-pass, sub-10ms latency
Ollaya – Ollama for open-source, Jev-style decision models
Ollaya is a local runtime for open-source decision models—think Ollama but for classification and scoring. It produces answers in a single forward pass with no token-by-token generation. A five-question request to the Laya model on an RTX 4090 takes about 8–10 ms end-to-end. Weights are pulled directly from Hugging Face, pinned to a commit and sha256-checked; the runtime uses ONNX Runtime and listens on 127.0.0.1 by default. It speaks TypeSafe's /v1/systemone API, so the TypeSafe Python SDK 0.7.1 works unchanged against a local server. Four model families are available: Laya, decider, nli, and gliclass, ranging from 322M to 1.9B parameters, covering English and 100+ languages. The post does not disclose training data sources or fine-tuning details. Desktop apps cover macOS, Windows, and Linux; a Docker image is also provided. GPU acceleration requires NVIDIA driver R580 or newer.
Why it matters: Ollaya packages classification/scoring models as an Ollama-style local tool — clean concept, solid latency data (8-10ms vs 200ms+ for hosted APIs). But it's early beta with no model ecosystem or real-world deployment stories yet, so it lands at the 72 featured threshold.