When your product is used by AI, not just humans: how to evaluate AI-friendliness
当产品不仅给人用,也给 AI 用:如何评估你产品的 AI 亲和度?
This piece argues that as AI coding tools become primary users of dev products, a product's AI-friendliness is a hard requirement. Supabase open-sourced supabase/evals and found that Agent failures often stem from unfriendly docs, CLI hints, or error messages—not model intelligence. Stripe, Convex, and Vercel are all building vertical regression evals instead of chasing generic leaderboards. Vercel's data is striking: default Agent Skills went unused in 56% of cases, yielding the same 53% pass rate as no docs; embedding an 8KB AGENTS.md index directly in context hit 100%. The post recommends pulling 20–50 real pain points from support tickets and GitHub Issues, then running a two-layer setup of static linting in CI plus dynamic sandbox evals. Fix the product side first on every failure before swapping models.
Why it matters: Fresh angle backed by concrete cases (Supabase, Vercel), not just theory. But it's a personal blog without primary data or exclusive interviews, so authority is limited—capped at 78, the featured threshold.