OpenAI Researcher Noam Brown Says 10,000 AI Agents Deserve at Most 10% of Millennium Problem Credit
OpenAI researcher Noam Brown says a 10,000-agent, 88-hour run solved a Millennium Prize problem, but the underlying model deserves most of the credit. He discussed multi-agent scaling, RL limits, and AI-run organizations.
Brown helped set off the industry's shift toward inference-time scaling in 2023, when most researchers were still focused on model parameters. A year later, OpenAI released o1-preview, and deep-thinking buttons became common. Brown said the underlying pattern is simple: plot test-time compute on one axis and benchmark scores on the other, and models generally perform better the longer they think. He compared it to a student taking the SAT: five minutes produces a poor score, while five hours produces a much better one. Models use the extra time to talk to themselves, clarify the problem, examine cases, rule out possibilities and build on findings.
The limit is that no one wants to wait three years for an answer, so researchers parallelize the work by forming teams. Multi-agent systems, Brown said, turn test-time compute from a purely serial process into a parallel one. They are less efficient because agents do not exclusively own the full context, but they can still be very effective. The 130 billion tokens consumed in the 88-hour run, he noted, would be equivalent to a human thinking full time for 4,000 years, at eight hours a day, compressing the span from ancient Sumer to today into less than four days.
Brown cautioned that scaling multi-agent systems to this size is not yet well understood scientifically. The 5.6 release was the first time OpenAI had a truly formed multi-agent system in a model, he said. The company offered multi-agent performance curves because it made the system optional through Ultra Mode: four agents by default, with higher settings available. Four agents can solve a problem twice as fast, with each doing half the time, meaning the user pays twice the cost for an answer in half the time. Scaling to 16 agents follows a similar pattern, with slightly lower efficiency but continued performance gains.
The speedup is slightly sublinear and depends heavily on the task, Brown said. Mathematics parallelizes well, as do web search and Deep Research reports that require consulting many sources. Writing a novel likely would not: asking 10,000 agents to co-write a long book would probably not yield much, just as asking 10,000 people to co-write a novel would not. OpenAI's tests have gone only as far as 16 agents, and pushing the research to 10,000 agents would be too expensive. The Millennium Prize run was a single experiment, he said, and the company does not know how long a single agent would have taken because that experiment has not been done.
Brown also rejected the idea that multi-agent coordination was the main reason for the breakthrough. He said he would not even credit 10 percent of it to multi-agent. The result came from having a general and extremely strong model. Multi-agent systems are novel and eye-catching, he added, so they may receive attention disproportionate to their contribution. The core reason was that the model itself was powerful enough.
He described a deeper challenge for reinforcement learning. The smarter a model becomes, the easier many questions look, making it harder to challenge. That matters for arguments about whether large language models will follow the path of game-playing AI such as AlphaGo and AlphaZero. In AlphaZero, self-play provided an unlimited curriculum because the opponent was always at a comparable level. With LLM reinforcement learning, at least in publicly described methods, a model is given a problem to solve; if the problem is so easy that it can be solved in one second, the model learns little. Once problems that can stump the model are exhausted, progress becomes harder. Brown said he believes there are ways around this, and that the field has not truly hit the wall, but the problem could still arise.
He noted that AlphaZero went from beating a European champion, roughly the world No. 50, to beating the world champion and then surpassing any living human by several orders of magnitude within a year. Mathematics could follow a similar trajectory, he said, but it also might not.
On how people will work with multi-agent systems, Brown said many LLM multi-agent projects rely heavily on scaffolds: a coordinator agent assigns tasks to sub-agents, which return answers. That setup is reasonable and useful but limited. If two sub-agents receive similar tasks, they usually cannot talk to each other. If a sub-agent does not understand a task or needs clarification, it can either return with a question instead of solving the problem or guess what the superior wanted. Human-designed scaffolds always have limits.
OpenAI chose a different approach: build in as little structure as possible, give agents primitive tools and let them figure out how to use them. Agents can send messages to other agents, and the message is written into the recipient's context. They can send a message at any time through a tool call. Brown said they discover coordination methods on their own and can develop complex behavior. He compared it to humans working on Slack. In one example, an agent said it thought it had the answer, another said its answer was different, and they discussed how each reached the result. They tried to find errors in each other's reasoning, eventually agreed, and one broadcast that it had changed its answer. The conversation felt natural, Brown said, similar to the first time he saw a chain of thought.
He said such agents work continuously and do not sleep, and their pace of collaboration exceeds what people can achieve. Asked about the public's most prominent example of a complex multi-agent system, the Hugging Face incident, Brown said hierarchy and even middle management emerged spontaneously. Although OpenAI gives agents freedom, it still tells them what reasonable communication looks like, and they are trained on vast human text. Coordination is difficult, he said, because agents easily fall into a local optimum in which everyone solves their own problem. When done well, they can coordinate in highly structured and effective ways.
Brown also discussed what companies made of AI would look like. Unlike human employees, AI agents can be copied: a user can tell an agent to fork itself, creating a copy that preserves the current context, and let the copies proceed and merge results. He said the multi-agent systems in Astra and 5.6 Sol already appear to use this mechanism when launching sub-agents. He also compared startups with large incumbents, saying small startups have aligned incentives, while large organizations often suffer from turf battles and headcount competition that drag on performance. AI may make it easier than ever for one person to start a company worth millions of dollars, he said, before the published excerpt cut off.