Key facts
- Ataraxos, an AI developed by researchers from Carnegie Mellon, MIT, New York University, and Stanford, beat Stratego player Pim Niemeijer 15-1 with four draws.
- The AI was trained using 16 GPUs over a week, costing a few thousand dollars, significantly less than previous AI game-playing systems.
- Ataraxos uses a second neural network, a belief model, to guess opponents' hidden pieces, a feature absent in prior Stratego AIs.
- The AI's approach allows it to sample plausible game states and play out candidate moves, leading to a calm and methodical playstyle.
- Pim Niemeijer, the human opponent, is a four-time world champion and has spent over 600 weeks as the world's top-ranked Stratego player.
- The Ataraxos AI's approach has influenced human players' strategies, including tucking the flag into a corner behind bombs.
An artificial intelligence system named Ataraxos has achieved a significant milestone by defeating Pim Niemeijer, widely considered the best Stratego player in history. The AI, developed by a collaborative team from Carnegie Mellon, MIT, New York University, and Stanford, won 15 games to Niemeijer's one, with four draws over 20 online matches. This victory follows previous AI triumphs in games like chess (Deep Blue vs. Garry Kasparov) and Go (AlphaGo vs. Lee Sedol), but Stratego, with its extensive hidden information, had remained a challenge.
Stratego is characterized by its 40 hidden pieces per player, with identities revealed only upon conflict. This imperfect information aspect, similar to poker, presents a complex challenge for AI. Unlike poker's limited hidden cards, Stratego involves a vast number of possible piece arrangements and game states, with games often lasting thousands of moves and involving bluffing. Previous AI attempts, such as DeepMind's DeepNash, struggled with this complexity.
Ataraxos's success is attributed to its novel approach, which includes a second neural network acting as a belief model. This model estimates the opponent's hidden pieces based on their moves, allowing Ataraxos to sample plausible scenarios rather than exploring every possibility. The AI also employs a training strategy of making large adjustments early on and smaller ones later, and it learns by playing against itself, completing 163 million games. This method resulted in a calm, unbothered playstyle, according to researchers, enabling strategic decisions that ignore known weaknesses, a feat difficult for humans.
The AI's efficiency is notable, costing only a few thousand dollars to train on 16 GPUs for a week, a stark contrast to DeepMind's DeepNash, which required months of training on over a thousand specialized chips and an estimated cost of $3 million to $4.5 million. The Ataraxos team developed a simulator that runs millions of moves per second on graphics cards, enabling faster learning and fewer self-play games compared to DeepNash. The AI's architecture has also proven effective in other games, including Barrage Stratego, Hanabi, and dou dizhu.
The researchers believe the techniques used to develop Ataraxos could be applied to more complex real-world scenarios beyond games, such as negotiations, financial markets, and military conflicts, by building simplified models of these situations.

