AREX 2, an agent trained to revise its own work across many rounds, with 27 billion parameters, released as an open weight model (its trained weights are downloadable) by the Beijing Academy of Artificial Intelligence (BAAI) and trained on long
An open-weight 27B agent from the AREX Team at BAAI keeps revising its own work across many rounds, and the habit transfers to four research benchmarks the agent never saw in training. The arXiv preprint argues the lever is not a bigger base model but a specific shape of training data: long revision trajectories, with the intermediate failures kept.
The paper, AREX-2, takes a 27B-class model from Alibaba's Qwen family and trains it on synthetic traces of an agent re-reading its work and trying again. Training tasks come from machine-learning engineering problems on MLE-bench Lite and algorithmic-programming problems on Frontier-CS, both domains where an answer can be checked.
The reported result is transfer. AREX-2 was tested on four benchmarks it never trained on: BrowseComp (web browsing), HLE (Humanity's Last Exam, a hard multidisciplinary question set), GAIA (assistant tasks needing browsing and documents), and DeepSearchQA (multi-step search). The paper reports 84.0, 52.6, 92.2, and 93.8 on those, plus 81.8 on MLE-bench Lite and 70.7 on Frontier-CS in training. The first-party repository documents 300 calls per round, 1,500 total, with an external judge.
The claim is narrow and testable. Self-improvement here means revising a solution across test-time rounds inside one task, not recursive modification of model weights. The paper does not show matched base-model or training-data ablations, and the improvement curve is bounded by round budget, not autonomous retraining. Whether the loop generalizes beyond the benchmark set remains open.