Ollaya runs open-source decision models locally with single-pass, sub-10ms latency
What happened
Ollaya 是一个本地运行开源决策模型的工具,可以理解成分类和打分版的 Ollama。它不走逐 token 生成,而是一次前向传播直接给出答案。在 RTX 4090 上跑 Laya 模型,处理一个包含五个问题的请求,端到端耗时约 8 到 10 毫秒。模型权重直接从 Hugging Face 拉取,锁定到具体 commit 并用 sha256 校验,运...
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pagePickOllaya runs open-source decision models locally with single-pass, sub-10ms latency
Ollaya is a local runtime for open-source decision models—think Ollama but for classification and scoring. It produces answers in a single forward pass with no token-by-token generation. A five-question request to the Laya model on an RTX 4090 takes about 8–10 ms end-to-end. Weights are pulled directly from Hugging Face, pinned to a commit and sha256-checked; the runtime uses ONNX Runtime and listens on 127.0.0.1 by default. It speaks TypeSafe's /v1/systemone API, so the TypeSafe Python SDK 0.7.1 works unchanged against a local server. Four model families are available: Laya, decider, nli, and gliclass, ranging from 322M to 1.9B parameters, covering English and 100+ languages. The post does not disclose training data sources or fine-tuning details. Desktop apps cover macOS, Windows, and Linux; a Docker image is also provided. GPU acceleration requires NVIDIA driver R580 or newer.