Skip to content
Hacker News front page

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Orca-Bench: How Ready Are Language Model Agents for Oncall?

Five frontier coding agents were tested on 1,079 production-style RCA tasks in a live microservice testbed with Prometheus, Jaeger, and OpenSearch. The best agent hit 25.3% root-cause accuracy on medium-difficulty tasks and 10% on hard ones. The weakest model hallucinated an implausible root cause in 40% of reports. Removing source-code access hurt every metric. The authors stress that real production systems are orders of magnitude larger, so these numbers are a lower bound—agents aren't ready for unsupervised oncall.

Why it matters: A production-realistic benchmark (Prometheus, Jaeger, OpenSearch, 6 days of microservice data) testing five top coding agents on 1,079 root cause analysis tasks. Results are brutal: 25.3% accuracy on medium, 10% on hard, and the worst model fabricated root causes in 40% of rep...

Read the original ↗Export Markdown