A third of 36 popular MCP servers fail agents on usability
I graded 36 popular MCP servers on agent usability. A third got a D or F
Teng Li linted 36 popular MCP servers with his tool mcpgrade: 11 scored D or F. The main culprit is undocumented parameters—firecrawl had 132 out of 134 errors from bare params, and MongoDB and Notion official servers are similarly bare. In live model evals, poorly documented servers dropped tool-selection accuracy from 100% to 84%, and refusal rate on out-of-scope tasks fell from 100% to 50%. The fix is simple: add .describe() to every parameter. context7 already did it and jumped from C to a perfect score.
Why it matters: A first-person experiment with a custom linting tool, quantifiable accuracy drops (100% → 84%), and named servers with specific failure modes. Hits all three HKR axes, but it's an engineering practice piece rather than a product launch or model breakthrough, so it lands in the...