On October 11, 2026, The Decoder covered a Vals AI study comparing GPT‑6 Sol and Claude Opus 5.5 as single agents versus multi‑agent teams on a web‑app benchmark. Teams cost between about 2 and 5 times more tokens, and only one of four setups showed a statistically significant quality improvement.
This article aggregates reporting from 2 news sources. The TL;DR is AI-generated from original reporting. Race to AGI's analysis provides editorial context on implications for AGI development.
The Vals AI study is a useful reality check on the current hype around agent swarms. On a realistic coding benchmark where models build and deploy full web apps, multi‑agent teams of GPT‑6 Sol and Claude Opus 5.5 delivered almost no statistically meaningful quality gains in three of four configurations, while burning two to five times as many tokens and, in Claude’s case, often taking far longer to finish. That suggests most of the value of today’s strongest models still comes from their single‑agent capabilities, not from elaborate orchestration layers. ([the-decoder.com](https://the-decoder.com/ai-agent-teams-waste-massive-tokens-for-barely-measurable-quality-gains-research-finds/))
This does not mean agent teams are pointless, but it pushes the burden of proof back onto people selling “1,000‑agent” architectures as a magic multiplier. If adding agents mostly buys speed at a steep cost, it is a performance trade‑off problem, not a path to emergent super‑capabilities. For frontier labs, the message is that algorithmic and training improvements remain the primary levers; you cannot simply stack more agents on top of a mediocre core model and expect AGI‑like behaviour.
For the AGI race, that may be good news. It implies that scaling intelligence is still mostly a function of model quality and compute, which are measurable and somewhat governable, rather than opaque multi‑agent dynamics that are harder to predict or regulate.