Skip to content
Hacker News front page

Manifest deprecated its LLM router, arguing the savings get spent elsewhere

Everyone is building LLM routers, we deprecated ours

Manifest launched an LLM router in March that classified requests into four complexity tiers to cut costs by picking cheaper models for simple tasks. After four months and 7,000 cloud users, they deprecated it in June and will shut it down September 1. The main problems: prompt text alone can't reveal true complexity—'evaluate the tests for $GIT_REPO' is trivial for a static site and brutal for the Linux kernel. Cache reads are 75–90% cheaper than uncached inputs, and prefix caching naturally makes the router stick to one model, defeating its own purpose. Switching models mid-session breaks consistency, makes tools harder to master, and adds uncertainty to evals and observability in automated workflows. Manifest's takeaway: for most use cases, a single battle-tested model beats routing—the money saved on inference gets paid back somewhere harder to measure.

Why it matters: Manifest's four-month production postmortem on their LLM router has concrete failure modes with real scenarios and numbers — not hand-waving. Score held back because it's a single-vendor anecdote with no controlled comparison, and the article body is truncated so the full argu...

Read the original ↗Export Markdown