AdvertisementAdvertisementAdvertisementAdvertisement
AI

Stanford Study: AI Teams Outperform Individual Models

9/25/2026, 03:48 PM • Evgenia Sliv

(edited: 09/25/2026)

Stanford Study: AI Teams Outperform Individual Models

A study from Stanford University has shown that self-organizing teams of artificial intelligence agents outperform both individual models and traditional multi-agent approaches based on debates and voting.

The new paper presents a methodology in which groups of AI agents learn to collaborate based on a small set of past tasks and then apply these teamwork strategies to completely new tasks. This approach is called SAT (Self-Organizing Teams). It achieved an average accuracy of 66.7% across five mathematical and physical tests – significantly higher than the best individual agent in the group (48.8%).

Traditional multi-agent systems typically follow a rigid scenario: agents discuss a problem and then vote on a solution. SAT employs a fundamentally different approach – teams of agents develop their own organizational structures, including roles, participation rules, dialogue stages, and information flow schemes.

The key idea is learning from surprisingly small datasets. SAT teams developed their collaborative guidelines using only 15 AIME 2024 tasks or 25 GPQA Diamond tasks, and then transferred these strategies unchanged to completely different test sets they had never encountered before. The concept is largely borrowed from organizational psychology. Agents exchange reasoning, challenge each other's logic, and combine partial solutions to arrive at answers unattainable by a single agent.

The list of models involved reads like a star-studded AI team. For the mathematics and physics tasks, o3-mini, Claude Sonnet 4, and DeepSeek-V3 were used, while Gemini-2.5-Flash, Llama-4-Maverick, and GPT-4.1 were employed in knowledge and logic tests.

Popular news