Echo routes prompts across open-weight models, claiming Fable-level results at 1/3 the cost
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
Echo is an experimental system from TracerML that pools open-weight models like GLM-5.2 and Kimi K2.7, then decides per request which models to invoke and how to combine their outputs. The author first computed a theoretical upper bound—if you could always pick the best model combination after seeing results, performance far exceeds any single model. Echo tries to approach that bound without knowing the answers in advance. On the author's own eval mix, Echo matched Fable's aggregate score at roughly one-third the inference cost. The post does not disclose specific benchmark names or absolute scores; methodology is at echo.tracerml.ai/eval. Known issues: routing and combination decisions sometimes fail, and the author is testing whether the approach holds for coding and agentic tasks. Caveat: the eval set is self-built, so saturation and representativeness are unknown—don't rush to benchmark against Fable yet.
Why it matters: Show HN project that dynamically routes across open-weight models (GLM-5.2, Kimi K2.7) to approximate Fable-level results at 1/3 cost. Concrete mechanism and numbers, but the eval set is self-built and the post doesn't disclose comparison details or sample size against Fable —...