$8,000 Beats Four-Time World Champion: AI System Ataraxos Ends Stratego's Last Human Bastion

A research team from Carnegie Mellon, MIT, New York University, and Stanford built Ataraxos, an AI system that defeated four-time Stratego world champion P

On September 30, 2026, a paper published in Nature announced with a set of concrete numbers the fall of the last human stronghold in board games: the AI system Ataraxos defeated four-time world champion Pim Niemeijer 15 wins, 1 loss, and 4 draws in a 20-game official series. This Dutch player has held the top spot on the Stratego world ranking list for more than 600 weeks and is widely recognized as the most decorated Stratego player in history. The joint team, composed of researchers from Carnegie Mellon, MIT, New York University, and Stanford University, accomplished the feat that DeepMind spent $3 million to $4.5 million but failed to achieve, with a training cost of less than $8,000, while also recording a 39-2 record in the Stratego World Championship.

Why Stratego Is Harder Than Chess

To understand the significance of this breakthrough, one must first understand the special challenge Stratego poses to AI. Chess and Go are perfect-information games—everything on the board is fully visible to both sides, and the optimal solution can in principle be approximated through sufficiently deep search. Stratego is another world: each side has 40 pieces representing different ranks. At the start, the opponent can see that your pieces exist but does not know their identities. Only when two pieces meet head-on does the weaker one leave the board, revealing its rank at the same time.

This hidden information produces exponential complexity. According to the paper's data, Stratego has more than 1033 opening setups, far exceeding Go's state space. A game of chess usually lasts around 40 moves, while Stratego can easily extend to 2,000 moves. Even harder is the strategic value of bluffing: a weak piece can disguise itself as the strongest piece to scare off an opponent, but if it bluffs too often, the opponent will see through it; if it never bluffs, its behavior becomes predictable. This "psychological warfare" at the game-theoretic level caused all previous AI methods to hit a bottleneck when scaled up.

Gabriele Farina, an assistant professor of electrical engineering and computer science at MIT and the paper's corresponding author, explained the scale of the problem:

"In Stratego, the possible universe you have to deal with explodes. AI techniques developed for games like poker absolutely cannot scale in this setting."

Three-Part Architecture: Self-Play, Belief Network, and Decision-Time Search

The technical core of Ataraxos is the coordination of three components. Individually, none is new, but integrating them in an imperfect-information setting is the paper's central contribution.

The first component is self-play reinforcement learning. The system accumulates playing strength through about 163 million self-play games; winning moves are reinforced, and losing moves are pruned. To prevent self-play from falling into cycles (for example, always sticking to a certain set of opening setups), the team introduced a regularization mechanism: early in training, the policy is forced to vary substantially, and as training progresses the magnitude of adjustments is tightened. This ensures diversity in the strategy repertoire and reproduces human players' intuition for maintaining "unreadability."

The second component is the belief network—the most critical innovation. Traditional search algorithms fail in imperfect-information settings because the state space explodes. Ataraxos's solution is to train an independent generative model specifically to infer the distribution of the opponent's hidden pieces from their movement traces. Rather than exhaustively enumerating all possibilities, it samples "the most credible board states" and reasons about candidate moves on those states. Farina describes it this way:

"We are not guessing blindly; we use decision-time planning to find the most credible board states. This generative model lets us genuinely focus on the specific board and opponent we face right now."

The third component is decision-time planning. The belief network provides a state distribution, the policy-value network scores each candidate move, and search finds the action with the highest expected payoff within a limited computational budget. The integration of these three allows Ataraxos at every step to replace "enumerating all possibilities" with "an informed guess," turning an unsolvable problem into a solvable one. The training process itself ran on 16 Nvidia H100 GPUs for about one week, plus an additional 4 days of training for the belief network, for a total computational cost below $8,000.

Comparison with the Precedent: What DeepNash's Failure Shows

DeepMind's DeepNash was the AI system previously closest to conquering Stratego, first unveiled in 2022. Its training scale far exceeded Ataraxos's: according to public information, DeepNash cost $3 million to $4.5 million and trained for about 5.5 billion games. Despite investing nearly 500 times Ataraxos's resources, DeepNash never reached a level beyond top human players in official matches.

One detail is even more noteworthy: the Ataraxos team proactively proposed a direct comparison test between the two systems, but DeepMind's response was that it could not be done—because DeepNash's code could no longer run. This fact indirectly reveals an engineering reality inside large AI labs: for projects that cost millions of dollars, without continuous maintenance, their engineering reproducibility may be far lower than outsiders expect.

From a historical sequence, Stratego is the latest in a series of "last bastions": Deep Blue defeated Kasparov in 1997, AlphaGo defeated Lee Sedol in 2016, and Texas hold'em bots have comprehensively dominated professional poker players in recent years. Each breakthrough was thought to require a new ability to "truly understand" information games, but each time, in retrospect, it was a product of engineering and algorithmic coordination. Ataraxos's difference lies in scale—it did not overwhelm the problem with more resources, but solved a similar problem with a smarter architecture at one percent of the cost.

The Real Significance for Real-World Tasks

The paper's authors themselves kept a cautious but specific phrasing regarding the technology's broader implications. Farina noted:

"In the imperfect-information tasks you face in reality, you often cannot hope to enumerate all possibilities. There are too many."
Ataraxos also achieved strong performance in three other imperfect-information games (Barrage Stratego, the cooperative card game Hanabi, and the Chinese card game Dou Dizhu), providing preliminary multitask evidence for the method's generalization ability.

For AI engineering practice, the more direct impact is a shift in the efficiency paradigm. If the combination of "belief network + decision-time search" can reproduce this efficiency advantage in other imperfect-information domains, then cybersecurity (guessing an adversary's intended actions), negotiation strategy (unknown opponent preferences), supply-chain games (hidden inventory information), and other scenarios will have a principled algorithmic path to follow, without relying on DeepNash-style brute-force compute stacking.

Samuel Sokota, a doctoral student at Carnegie Mellon and the paper's first author, contrasted the fundamental difference between Stratego and chess:

"This is completely different from chess—in chess, no matter how many times you play, the best move is still the best move."
This sentence highlights the core characteristic of imperfect-information games: strategies must be probabilistic and adaptive, because the optimal action depends on the opponent's hidden state, and that hidden state updates with the opponent's behavior at every step.