Skip to content
Computing Life · Share · Yage

AgentFlow trains a 7B decision node in the loop, gaining 17.2 points over swapping in GPT-4o

换 GPT-4o 只多 5.8 分,训练 7B 却多了 17.2 分

Stanford's AgentFlow paper shows that in the same agent orchestration, swapping a frozen Qwen2.5-7B decision node for GPT-4o adds only 5.8 points on average across six benchmarks. Training that same 7B node with real tool feedback adds 17.2 points. Only the Planner's selection policy is updated; the system skeleton stays fixed. The model learned to prefer Wikipedia over Google for medical queries, and tool-calling errors dropped by up to 28.4%. The post also lists four gates for real-world adoption: high-frequency tasks, automatic success verification, bottlenecks truly in decision logic, and a resettable environment. The cost story is incomplete—the paper discloses 8×A100 but not total training time or the cumulative bill for the GPT-4o judge.

Why it matters: AgentFlow from Stanford answers a concrete bottleneck question for agent builders: swapping in GPT-4o only adds 5.8 points, but training the 7B decision node on real execution feedback adds 17.2. Has numbers, mechanism, and engineering reproducibility—directly actionable signa...

Read the original ↗Export Markdown