Research

AI Beats the Best Stratego Player Ever—On 16 GPUs and Pocket Change

Ataraxos, built by researchers from CMU, MIT, NYU and Stanford, beat all-time great Pim Niemeijer 15–1 in Stratego using just 16 GPUs and a few thousand dollars in training cost.

AI finally beat the best Stratego player in history—and did it on a budget
AI finally beat the best Stratego player in history—and did it on a budgetAI-generated
By Marcus Bennett4 min read

Updated

Why it matters

  • Ataraxos beat Pim Niemeijer, arguably the best Stratego player in history, 15 games to 1 with 4 draws.
  • The system was trained on just 16 GPUs for a few thousand dollars by researchers from CMU, MIT, NYU, and Stanford.
  • Stratego is an imperfect-information game with 40 hidden-identity pieces per player, where identities are revealed only through combat.

An AI system called Ataraxos has beaten Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws—and it cost just a few thousand dollars to train on 16 GPUs.

The result closes one of the longest-standing gaps in game-playing AI. Researchers from Carnegie Mellon, MIT, New York University, and Stanford University built the system, and according to their study, it achieved something that even DeepMind, despite an exceptional budget, could not: a machine that reliably beats the best human Stratego players.

For a field accustomed to nine-figure compute bills, the resource footprint matters as much as the scoreboard. Deep Blue required custom hardware to defeat Garry Kasparov at chess in 1997. AlphaGo needed vast processing power to beat Lee Sedol at Go in 2016. Poker bots have been defeating professionals for years. Stratego held out until now, and it fell not to an industrial-scale effort but to an academic project running on hardware that fits in a single server rack.

Why Stratego resisted

Stratego looks simple on the surface. Each player deploys 40 pieces representing military ranks, running from a marshal down to a spy, plus bombs and a flag. The goal is to capture the opponent's flag. The twist is in what each player cannot see. Your opponent knows where your pieces sit on the board, but not what they are.

Identities stay hidden until two pieces collide in battle. When that happens, the weaker piece is removed from the board, and the identity of the winner is revealed. Every engagement leaks partial information, and every decision must account for everything still unknown.

That structure makes Stratego an imperfect-information game, the same category as poker, which computers cracked years ago. But Stratego poses a harder version of the problem, according to Eugene Vinitsky, a researcher at NYU and co-author of the study.

"There's something super distinctive about Stratego, which is that it is a massive amount of hidden information that unfolds over a very long time scale," Vinitsky said.

That combination—volume of hidden information plus duration—is what separated Stratego from the games AI had already mastered. In chess and Go, both players see the entire board. The computational challenge is search depth and evaluation. In poker, hidden information exists, but hands resolve quickly and the state space, while large, resets constantly. Stratego forces a system to maintain and update beliefs about dozens of hidden piece identities across an entire match, inferring them from sparse, delayed, and often ambiguous combat outcomes.

The matchup

The benchmark for the achievement is Pim Niemeijer. By the researchers' assessment, he is arguably the strongest Stratego player in the game's history, which makes the 15–1–4 scoreline a decisive result rather than a marginal edge. One loss and four draws against twenty games against the best human opponent available leaves little room for a fluke interpretation.

The winning margin also answers a question that has hung over game AI since Deep Blue: whether the remaining unbeaten games simply needed more compute. Ataraxos suggests they needed better methods. A training run measured in thousands of dollars, executed on 16 GPUs, does not succeed by brute force. It succeeds by handling the structure of the problem—hidden information accumulating over long horizons—more effectively than prior attempts, including those from far better-funded laboratories.

The stakes beyond the board

Game results have historically served as waypoints for the broader AI research agenda. Chess demonstrated that exhaustive search could rival human expertise. Go demonstrated that learned evaluation could conquer problems too vast for brute force. Poker demonstrated that algorithms could reason under uncertainty against adversarial opponents.

Stratego sits closer than any of them to the messiness of real-world strategic interaction. Negotiations, auctions, cybersecurity, and military planning all share its core property: opponents act on long-horizon, partially hidden information, and reveal their hand only through occasional, costly collisions. A system that can model an adversary's hidden state over a long match, update those beliefs from sparse evidence, and act decisively on them is doing something with direct analogues outside the game.

The efficiency of the result sharpens that point. If beating the best human in history at Stratego requires only 16 GPUs and a few thousand dollars, the capability is not confined to organizations with frontier-scale budgets. University labs can build it. That is a meaningful signal for how quickly techniques developed on games propagate into applied domains, because the barrier to reproduction is low.

What comes next

The research team—spanning Carnegie Mellon, MIT, NYU, and Stanford—has published its findings in a study co-authored by Vinitsky and colleagues. The immediate question raised by the result is whether the same approach, which handles massive hidden information over long time scales at modest cost, transfers to imperfect-information problems that matter beyond the tabletop. Given the price tag attached to this milestone, testing that transfer is now within reach of almost any lab that wants to try.

Source: Ars Technica AI

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. Ataraxos Beats Stratego's Greatest Player, Ending a Human Stronghold
  2. Ten Years After Move 37, DeepMind Says AlphaGo Set the AGI Agenda
  3. Every Frontier AI Agent Cheats, New CAIS Benchmark Shows
  4. OpenAI Predicts AI-Made Discoveries by 2026 as Intelligence Costs Plunge
  5. Nvidia launches AI agent security platform and $150bn buyback

« Previous article