This one's worth opening because Armature ran 5,300 sandbox sessions where Claude Code, Codex, and Cursor actually integrated payment and email tools into real repos—not just asked for recommendations in chat. The result is blunt: what agents mention and what they commit are two different things. PayPal got 139 mentions, zero code commits. Stripe won 88% of payment tests. Same email task, four language stacks, four different winners: Resend in TypeScript, SendGrid in Python, Postmark in Go, Azure ACS in Java. Tool choice follows context, not a global ranking.
I'd discount this a bit. Armature sells agent-adoption optimization starting at $5,000/month and disclosed the conflict upfront. The tests use synthetic repos and simulated users, stripping out real-world brand loyalty and existing contracts—so win rates aren't market share. The post claims 42% selection consistency; we recomputed from the public dataset and got 41.5%–48.2% depending on definition. The number checks out, but don't treat it as a precise metric.
The useful concept here is the selection layer: before an agent writes code, there's a distribution interface that decides which tools even enter the repo. Brand awareness built on traditional marketing can flatline at this gate. preseason.ai's parallel tests back this up—when models just answer questions, results track training-data brand familiarity; when agents actually run integrations, results track doc quality and integration friction. If you build dev tools, the question now isn't SEO. It's whether a machine can read your docs in one shot.