Kimi K3 ranks second on AA-Briefcase agentic benchmark, but costs more than Opus 4.8
Kimi K3: second only to Fable 5 on AA-Briefcase
Moonshot AI's Kimi K3 (2.8T params) scores 1543 Elo on AA-Briefcase, second only to Claude Fable 5 (1574) and a +727 jump over K2.6. Its rubric pass rate and analytical quality rival Fable 5, but presentation quality lags behind GPT-5.6 Sol and Opus 4.8. The catch: it averages 56 minutes and $10.57 per task—2.5x slower than Fable 5 and pricier than Opus 4.8. That's driven by 83 turns and 120k output tokens per task at $3/$15 per 1M input/output tokens. The post doesn't spell out which real-world workflows can tolerate that cost and latency.
Why it matters: Moonshot AI's Kimi K3 hits 1543 Elo on AA-Briefcase, second only to Claude Fable 5 at 1574, with a +727 generational leap — the closest a Chinese model has come to the top on an agentic knowledge-work benchmark. Score stays at 82 rather than higher because this is a third-part...