Is it agentic enough? Benchmarking open models on your own tooling
AI 智能体够格吗?在自有工具上评测开源模型
Hugging Face used its own transformers library as a testbed to see how much work open models really need when writing code, calling APIs, and debugging themselves. Instead of just scoring right or wrong, they measured how many detours and tokens each run took, and found that library docs and API design directly affect agent cost and success rate. The post lays out a full open-model harness with the pi coding agent but does not disclose final model rankings.
Why it matters: Hugging Face benchmarks open models' agent coding on their own transformers library, breaking down token costs and success rates — real engineering, not just scores. Hits all three HKR axes, but it's a methodology innovation rather than a capability breakthrough, landing at 78...