Skip to content
Trending storyDeveloping

Vals AI tests multi-agent teams: 1.8x to 5.1x the cost, limited quality gain

1 report1 sourceupdated 2 hours ago

What happened

AI digest

On October 11, 2026, The Decoder reported on a Vals AI evaluation of the cost and quality trade-offs of multi-agent teams. Vals AI tested GPT-6 Sol and Claude Opus 5.5 on Vibe Code Bench, comparing agent teams against single agents. Teams cost 1.8 to 5.1 times as much as a single agent, with limited quality improvement. The report gives no per-model quality scores or cost breakdown.

Written by AI from the coverage · updated 2 hours ago

Coverage

Follow the reports to see the story from different sides.

Oct 11
  1. The Decoder
    AI agent teams waste massive tokens for barely measurable quality gains, research finds

    Vals AI 在 Vibe Code Bench 上测试 GPT-6 Sol 和 Claude Opus 5.5,智能体团队成本为单智能体的 1.8x 至 5.1x,质量收益有限。

Heat over time

Not enough continuous observations to draw a trend yet.