AI Finally Conquers Stratego: The Last Major Game Frontier
Fifteen wins. One loss. Four draws.
That is the final tally from a match that closed one of the longest-running chapters in game AI. A system called Ataraxos, built by researchers from Carnegie Mellon, MIT, New York University, and Stanford, defeated Pim Niemeijer—widely regarded as the strongest Stratego player in history—by a margin that leaves little room for argument. The result, reported in October 2026, marks the first time a machine has reliably outplayed elite human competition at a game that had resisted every previous attempt.
The lineage here matters. Deep Blue toppled Garry Kasparov at chess in 1997. AlphaGo beat Lee Sedol at Go in 2016. Poker bots have been beating professionals for years. Each of those milestones fell in sequence, and each seemed to confirm a pattern: given enough compute and enough cleverness, any game would eventually yield. Stratego was the exception that kept proving the rule wrong. Even DeepMind, with resources that dwarf most academic labs, could not build a machine that consistently beat the best human players.
Now that barrier is gone—and it fell cheaply. Ataraxos trained on just 16 GPUs and a few thousand dollars. The contrast with the scale of prior efforts is the story inside the story.
What Makes Stratego So Difficult for Artificial Intelligence
Place 40 pieces on a board. Each one represents a military rank, from a marshal at the top down to a spy, plus bombs and a flag. Your opponent can see where every piece stands but has no idea what any of them are. You win by capturing the flag. Identities surface only when two pieces collide—the weaker one dies, and the survivor's rank is exposed.
Read next Laika's Wildwood: Stop-Motion Fantasy at TIFF 2026That single design choice turns Stratego into an imperfect-information game. Both players act with only partial knowledge of the true state of the board, and that hidden information is enormous in scope. As the researchers put it, Stratego involves "a massive amount of hidden info," and that phrase carries real weight when you try to formalize the problem.
Compare it to poker, which computers cracked years ago. Poker hides a handful of cards per player. The uncertainty is real but bounded—you can enumerate the plausible hands an opponent might hold, assign probabilities, and reason about bluffing frequencies. Stratego hides the identity of nearly an entire army across dozens of pieces, and those identities interact in ways that compound. A bomb positioned in one square changes the value of every piece that might approach it. A spy sitting quietly can erase a marshal. The belief space—the set of all board configurations consistent with what you have observed—explodes.
This is why Stratego AI proved so stubborn. Algorithms that thrive on perfect information, like the search methods that powered chess engines, have no direct answer for a game where you cannot see the position. And methods built for poker's smaller uncertainty struggle when the hidden state is this large and this structured. The game sits in an awkward middle ground: too much hidden information for brute-force search, too much spatial and tactical complexity for the abstractions that made poker tractable.
Meet Ataraxos: The AI That Beat the Best Stratego Player Alive
The team behind Ataraxos spans four institutions: Carnegie Mellon, MIT, NYU, and Stanford. That academic pedigree matters for how the result should be read. This was not a lone experiment or a publicity stunt. It was a coordinated research effort across some of the strongest computer science programs in the world, and the output was a system that beat the best human alive by a score of 15-1-4.
The efficiency is the headline for anyone who follows AI training costs. Sixteen GPUs. A few thousand dollars. DeepMind's earlier attempts at Stratego, by contrast, came from an organization whose budgets and compute reserves are measured on an entirely different scale—and those attempts did not produce a machine that could beat top humans reliably. The summary does not disclose exactly why those efforts fell short, and it would be speculation to say. What can be said is that Ataraxos achieved the milestone at a fraction of the cost, which points to algorithmic insight rather than raw horsepower as the deciding factor.
That distinction has become a recurring theme in AI research. When a problem falls to brute force, the result tells you about the state of hardware. When it falls to a modest cluster of 16 GPUs, the result tells you about the state of ideas. Ataraxos belongs in the second category.
The Match: Ataraxos vs. Pim Niemeijer
Pim Niemeijer is not a casual opponent. He is described as arguably the best Stratego player of all time, which makes him the correct benchmark for a claim of this magnitude. Beating a strong player in a few games proves little; variance and luck can carry a weaker opponent through a short series. Beating the best player in history across 20 games, with 15 wins against a single loss and four draws, is a different kind of evidence.
The score line also says something about the nature of the win. Four draws in a 20-game match suggest that Niemeijer was not simply overwhelmed—there were positions he held. One loss confirms that Ataraxos is not infallible. This is not a machine that solved Stratego in the sense of playing perfectly. It is a machine that plays better than any human, consistently, and that is the standard that matters for a game of this complexity.
For context on how rare this is: chess fell in 1997, Go in 2016, and poker over a span of years as bots steadily outperformed professionals. Stratego outlasted all of them. Niemeijer had the misfortune of being the human standing at the frontier when it finally gave way.
Broader Implications for AI and Real-World Decision Making
Stratego is a game, but the structure of its difficulty mirrors problems that matter outside the board. Hidden information is the norm in business, medicine, diplomacy, and military planning. You rarely know the full state of the world. You act on partial observation, watch for revealing signals, and update your beliefs as new information arrives—often at a cost, because in Stratego, as in life, the act of probing for information can destroy the piece doing the probing.
Poker research already demonstrated that AI can handle bluffing, deception, and probabilistic reasoning at a high level. Stratego adds spatial structure, long tactical sequences, and a belief space orders of magnitude larger. A system that handles all three at once, and does so after training on 16 GPUs, suggests that the techniques involved may generalize to domains with similarly sprawling uncertainty.
The cost point deserves emphasis. Real-world decision systems are constrained by budgets. If the methods behind Ataraxos scale down as well as this result suggests, they become relevant to organizations that could never afford a DeepMind-scale training run. That is a meaningful shift in who gets to build capable AI.
What Comes Next for Imperfect-Information AI Research
Every game milestone raises the same question: what is left? Chess fell, then Go, then poker, and now Stratego. The pattern has been remarkably consistent—each solved game becomes a testbed, and the techniques migrate outward.
For imperfect-information research, Stratego is a particularly useful testbed because its uncertainty is large and structured rather than small and enumerable. Methods that succeed here are more likely to transfer to messy real-world settings than methods tuned to poker's narrower hidden state. Expect the Ataraxos approach to be probed, extended, and stress-tested against variants of the game and against other domains with heavy hidden information.
What the result does not establish is a timeline for the next milestone. It took nearly a decade between Deep Blue and AlphaGo, and years more to reach this point. Progress in this field arrives in unpredictable jumps, often when a research team finds the right abstraction rather than the right amount of compute.
What is clear is that the list of games considered out of reach for AI just got shorter—and the machine that crossed this one off cost less than a used car.
Source: Ars Technica - All content



