Skip to content
Hacker News front page

Kimi K3 ranks second on AA-Briefcase agentic benchmark, but costs more than Opus 4.8

Kimi K3: second only to Fable 5 on AA-Briefcase

Moonshot AI's Kimi K3 (2.8T params) scores 1543 Elo on AA-Briefcase, second only to Claude Fable 5 (1574) and a +727 jump over K2.6. Its rubric pass rate and analytical quality rival Fable 5, but presentation quality lags behind GPT-5.6 Sol and Opus 4.8. The catch: it averages 56 minutes and $10.57 per task—2.5x slower than Fable 5 and pricier than Opus 4.8. That's driven by 83 turns and 120k output tokens per task at $3/$15 per 1M input/output tokens. The post doesn't spell out which real-world workflows can tolerate that cost and latency.

Why it matters: Moonshot AI's Kimi K3 hits 1543 Elo on AA-Briefcase, second only to Claude Fable 5 at 1574, with a +727 generational leap — the closest a Chinese model has come to the top on an agentic knowledge-work benchmark. Score stays at 82 rather than higher because this is a third-part...

Read the original ↗Export Markdown