Skip to content
Trending storyPast story

Ollaya runs open-source decision models locally with single-pass, sub-10ms latency

1 report1 sourceupdated 4 days ago

What happened

Summary

Ollaya 是一个本地运行开源决策模型的工具,可以理解成分类和打分版的 Ollama。它不走逐 token 生成,而是一次前向传播直接给出答案。在 RTX 4090 上跑 Laya 模型,处理一个包含五个问题的请求,端到端耗时约 8 到 10 毫秒。模型权重直接从 Hugging Face 拉取,锁定到具体 commit 并用 sha256 校验,运...

Coverage

Follow the reports to see the story from different sides.

Sep 26
  1. Hacker News front pagePick
    Ollaya runs open-source decision models locally with single-pass, sub-10ms latency

    Ollaya is a local runtime for open-source decision models—think Ollama but for classification and scoring. It produces answers in a single forward pass with no token-by-token generation. A five-question request to the Laya model on an RTX 4090 takes about 8–10 ms end-to-end. Weights are pulled directly from Hugging Face, pinned to a commit and sha256-checked; the runtime uses ONNX Runtime and listens on 127.0.0.1 by default. It speaks TypeSafe's /v1/systemone API, so the TypeSafe Python SDK 0.7.1 works unchanged against a local server. Four model families are available: Laya, decider, nli, and gliclass, ranging from 322M to 1.9B parameters, covering English and 100+ languages. The post does not disclose training data sources or fine-tuning details. Desktop apps cover macOS, Windows, and Linux; a Docker image is also provided. GPU acceleration requires NVIDIA driver R580 or newer.