An AI called Ataraxos used roughly 1/500th of prior work's compute to beat a four time world champion, a result the team attributes to the algorithm rather than the hardware.
A small academic team beat arguably the greatest Stratego player in history, 15 wins to one with four draws, on roughly 16 GPUs and what the researchers describe as a few thousand dollars of training compute. The match ran one fixed AI strategy through 20 games against Pim Niemeijer, a four-time world champion, and the result was reported in Nature on Wednesday.
Stratego is a two-player strategy game in which every piece starts hidden and the winner is the one who captures the opponent's flag. The hidden setup is the part that broke earlier systems. Where chess and Go let a program see everything, Stratego forces it to reason about an opponent's position it cannot observe. The MIT team and their collaborators at Carnegie Mellon, NYU, and Stanford put the search space at more than a decillion possible board configurations per game, a scale that punishes brute-force search. DeepMind's 2022 system, DeepNash, reached a strong amateur-human level but did not get to elite level; Ataraxos is the first reported AI to do so, according to the project site and the MIT news release.
The efficiency comes from how the system reasons at decision time, not from a larger training run. Before each move, Ataraxos samples possible hidden states from a belief network that has been trained on the system's own final policies. It then evaluates candidate moves through depth-limited rollouts and updates the policy with a single step of tabular magnetic mirror descent, a refinement step the authors adapted from the optimization literature. The belief network and the move policy are separate transformer components, trained in sequence, and only the move-policy self-play is the expensive piece. The paper reports 163 million finished games in one week of reinforcement learning on 16 H100 GPUs, plus a separate four-day run to train the belief network on four H100s.
The authors compare their reinforcement-learning run against prior work and report roughly 1/500th of the compute cost, 1/30th of the self-play games, and 1/100th of the training examples. Those ratios describe a single training run, not an audited all-in research-and-development figure, and the project has not published dollar costs. The relevant comparison is methodological: the same hardware budget that fails at chess-era approaches turns out to be enough when the algorithm does the work.
That is the access argument. The capability was previously associated with a frontier lab; the Ataraxos paper and the Ars Technica report show four university groups producing an elite result on hardware a research lab can rent. MIT frames future work around auditability, not deployment, and the documents say nothing about commercial use, military application, or operational transfer. The reported efficiency is a research result, not a product.
Two caveats belong in the same breath as the headline. The human benchmark was a single elite player against a fixed AI strategy; the team is not claiming a tournament field result. No independent reproduction has been reported, the Nature extraction was truncated in this review, and the supplementary materials have not been checked. The team is calling this a breakthrough, and the field has not yet verified it.
The next watch items are concrete: whether the same sample-believe-refine loop generalizes to other imperfect-information contests, and whether the 1/500th compute ratio holds up in independent replications.