Skip to content
Hacker News front page

Ramp launches a model router that claims to cut inference costs by 40% on average

Ramp Launches a Model Router

Ramp applies its cost-cutting DNA to model inference. Router is a single-endpoint gateway that picks the cheapest model meeting your performance bar per request, covering Anthropic, OpenAI, Grok, Fireworks, and others. One demo shows a $45.62 Router run vs. $297.85 for a generic frontier model. Customer Delphi reports a 92% model cost drop after running billions of tokens through it. Routing is free through 2026 with $26 in credits. The post doesn't disclose routing latency, fallback logic, or independent benchmarks.

Why it matters: Ramp launches a model router that auto-picks the cheapest model meeting your performance needs, with a demo showing costs dropping from $298 to $45. Directly relevant for teams running heavy inference, but it's a fresh launch with no third-party benchmarks yet, so the score st...

Read the original ↗Export Markdown