Ataraxos AI Beats Stratego World Champion, Findings Published in Nature
AI Ataraxos beat Stratego champion Pim Niemeijer 15-1 with four draws, using a belief model and efficient self-play.
Stratego is a game of hidden information: players cannot see the identities of enemy pieces. Ataraxos's main advantage was a second neural network that estimates those identities, allowing the system to search plausible game states rather than brute-force a massive hidden-information space. Like DeepMind's DeepNash, introduced in 2022, Ataraxos learned by playing against itself, for a total of 163 million games. Moves that led to wins were reinforced and played more often in future matches; moves that led to losses were played less. The difference was in how much Ataraxos adjusted after each game, because hidden information tends to send self-play learning algorithms around in circles. The team addressed this by making big, bold strategy changes early in training and small, careful ones later.
The bigger innovation was thinking ahead before each move. AIs such as AlphaGo refine their general strategy with a search just before acting. DeepMind could not make that work in Stratego because the search space was too large, leaving it an open question whether it was worth trying. "This is one of the things that we did figure out how to do," Farina said. The solution was a belief model, a second neural network trained to guess the opponent's hidden pieces based on how they had been moving. Instead of iterating through every possible arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on how they turn out.
Ataraxos was also far cheaper to train than DeepNash. DeepNash was trained for two to three months on 1,024 of Google's specialized chips, a run the Ataraxos team estimates would cost $3 million to $4.5 million at 2025 prices. Ataraxos needed 16 GPUs for a week, plus an additional four GPUs for four days to train the belief model. Farina and lead author Samuel Sokota achieved that efficiency by writing a simulator that runs millions of moves per second on graphics cards. "At the scale that we are in academia, we don't really have access to an entire field of GPUs," Farina said. The algorithm also learned faster: it played about 34 times fewer games than DeepNash and still ended up much stronger.
Editor's Summary Ataraxos, built by researchers at Carnegie Mellon, MIT, NYU and Stanford, beat four-time Stratego world champion Pim Niemeijer 15-1 with four draws. The system used a belief model to estimate hidden enemy pieces and a self-play method that trained on 163 million games with far less computing power than DeepMind's DeepNash. Its results were published in Nature.