AI Finally Conquers Stratego: The Last Great Hidden-Information Game
In 1997, Deep Blue took down Garry Kasparov at chess. In 2016, AlphaGo beat Lee Sedol at Go. Poker bots have been beating professionals for years. One classic game held out: Stratego. Even DeepMind, with its exceptional budget, couldn't reliably build a machine that beat the best human players. Now a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer — arguably the best Stratego player of all time — 15 games to one, with four draws. Training it took just 16 GPUs and a few thousand dollars.
That last detail may be the most striking part of the story. For three decades, each major game milestone has been defined as much by its cost as by its cleverness. Chess fell to specialized hardware, Go to a sprawling distributed training run, and heads-up poker to years of focused effort from well-funded labs. Stratego, a board game most people learn in childhood, quietly resisted all of that. Ataraxos suggests the barrier was never raw compute. It was the structure of the problem.
Why Stratego Resisted AI for So Long
Stratego hands each player 40 pieces representing military ranks, from a marshal down to a spy, plus bombs and a flag. You win by capturing the opponent's flag. Your opponent can see where your pieces sit, but not what they are. Identities surface only when two pieces collide in battle: the weaker one is removed, and the winner's identity is revealed.
Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026That single mechanic creates a game with more in common with poker than with chess. Chess and Go are games of complete information. Both players see the full board, so the core challenge is search and evaluation. Poker and Stratego are imperfect-information games. You must reason about what your opponent's hidden pieces probably are, update those beliefs as evidence accumulates, and act on probabilities rather than certainties.
"There's something super distinctive about Stratego, which is that it is a massive amount of hidden info," the research team noted in describing the problem. The scale is what separates it from poker. A poker hand involves a handful of hidden cards. A Stratego position involves 40 hidden identities whose locations are known but whose values are not, and where every collision leaks information to both sides. The belief space is enormous, and it shifts with every move, not just every showdown.
That combination — deep hidden information layered on top of a large spatial board — is why the game outlasted efforts from some of the best-resourced AI labs in the world. DeepMind's methods, so effective on Go, did not translate into reliable dominance over elite human Stratego play. The game required something different.
Meet Ataraxos: The AI That Cracked the Code
Ataraxos emerged from a collaboration spanning four institutions: Carnegie Mellon, MIT, NYU, and Stanford. That academic footprint matters. It signals a problem tackled by a distributed research team rather than a single corporate lab with a nine-figure compute budget.
The headline number is 16 GPUs. Not 16,000. Not a dedicated data center. Sixteen graphics processors and a training bill in the low thousands of dollars. For context, the compute behind AlphaGo's victory over Lee Sedol was orders of magnitude larger, and the engineering effort behind professional-grade poker bots consumed years of sustained investment from well-funded organizations.
To be fair, training cost is not the same as total cost. Research teams still spend on salaries, experimentation, and the false starts that never make it into a paper. The 16-GPU figure describes the final successful run, not the whole journey. Still, the gap is large enough to be meaningful. A result that once looked like it required an industrial-scale lab was produced with resources available to a well-equipped university group.
The pattern echoes a broader trend in AI research. As methods mature, the compute needed to reproduce frontier results tends to fall. What was a supercomputer problem in one era becomes a workstation problem in the next. Ataraxos is a clean example of that dynamic applied to a domain that had resisted previous approaches.
The Match: Ataraxos vs. Pim Niemeijer
The result was not close. Across 20 games, Ataraxos won 15, lost one, and drew four. That is a 75 percent win rate against the player widely regarded as the strongest in the history of the game, with a single defeat across the entire match.
For comparison, consider how the earlier milestones actually looked. Deep Blue beat Kasparov 3.5–2.5 in their 1997 rematch — a narrow margin that could have swung the other way. AlphaGo lost the fourth game of its 2016 series against Lee Sedol, a defeat that became a cultural moment in its own right. Poker bots have likewise won by margins that left room for human counterplay.
A 15–1–4 scoreline leaves very little room for that narrative. The draws matter too. In an imperfect-information game, draws can reflect genuinely balanced positions rather than cautious play, and four of them against an elite opponent suggests Ataraxos was not simply exploiting a single weakness. It was winning across the range of situations the game produces.
What This Means for AI and Imperfect-Information Research
Stratego occupies a specific niche that makes this result more than a novelty. It sits between two poles. At one end are complete-information games like chess and Go, where the main technical challenge is search. At the other end are games like poker, where hidden information is the entire problem but the state space is relatively compact.
Stratego combines both difficulties. Players face a large board with many pieces, a long horizon of moves, and pervasive uncertainty about what each opposing piece actually is. Solving it requires maintaining and updating beliefs about dozens of hidden variables simultaneously while still planning ahead. Methods built for one pole do not automatically work at the other.
That makes Ataraxos a useful data point for anyone working on real-world problems with the same shape. Military planning, negotiation, cybersecurity, supply chain disruption, and financial markets all involve hidden intentions, partial observations, and adversaries who adapt. None of them are board games, and no one should overstate the transfer. But the technical machinery for reasoning under deep uncertainty is shared, and progress on Stratego narrows the gap between game AI and messier domains.
The result also reframes how the field should think about difficulty. Stratego was not hard because it needed more compute. It was hard because it needed the right treatment of hidden information at scale. Once that was solved, the compute requirement turned out to be modest.
The Democratization of Advanced AI Research
Sixteen GPUs and a few thousand dollars. Those numbers carry their own message, separate from the 15–1–4 scoreline. They suggest that the frontier of game-solving AI is no longer reserved for labs with the largest budgets.
That shift has consequences. When a result like this is reachable by a multi-institution academic team, more groups can attempt similar problems. More attempts mean faster iteration, more diverse approaches, and a broader pool of people who understand the methods. The institutions behind Ataraxos — Carnegie Mellon, MIT, NYU, and Stanford — are elite, but they are not DeepMind. Their success demonstrates that the gap between well-funded corporate labs and academic research is narrower than the compute arms race sometimes implies.
There is a caution buried in the good news. If a few thousand dollars of training can produce a system that dominates the best human in a complex hidden-information game, that capability is not going to stay confined to a handful of organizations. The same accessibility that accelerates legitimate research lowers the barrier for anyone else. Stratego itself is harmless. The methods are not inherently so.
For now, the significance is simpler. After chess, Go, and poker fell, Stratego was the holdout — the game where hidden information and board complexity combined to defeat the best efforts of the best-funded labs. A team of academics with 16 GPUs just closed that chapter. The score was 15–1–4, and the budget was smaller than most people would guess.
Source: Ars Technica - All content



