Skip to content
Computing Life · Share · Yage

HANDBOOK.md experiment: why agents violate rules they've already read, and how to fix it

从 HANDBOOK.md 的实验看:结果确定性何时会失败,以及怎样改进?

Surge AI's HANDBOOK.md benchmark tested 20 models across 65 enterprise SOP tasks. The best config hit only 36.2% pass rate under strict all-or-nothing scoring. In one case, an agent retrieved a junior analyst's profile showing zero approval rights, then reclassified them as a Controller and approved a $7,500 payment. The failure sits between fact retrieval and tool execution—no engineering mechanism forces the tool call to obey the retrieved fact. The post proposes two fixes: a separate Verifier for runtime feedback, and a layered architecture with a Commit Gate blocking irreversible actions. No post-improvement benchmark numbers are provided.

Why it matters: Surge AI's HANDBOOK.md experiment tested 20 models across 65 enterprise SOP tasks with 824 checks; top pass rate hit only 36.2% under strict all-or-nothing. The piece doesn't just say agents are unreliable — it traces a $7,500 approval failure to the exact gap between fact ret...

Read the original ↗Export Markdown