The unit of account on agent leaderboards is structurally biased toward mean accuracy. Vendors report the average across runs, then readers treat that number as if it measures whether the same task succeeds again. It does not. The metric that production actually depends on is whether the agent clears the same task on every back-to-back try, and almost no public scoreboard surfaces it.
IBM Research just made that gap measurable. On AppWorld, a common step-by-step agent using GPT-4.1 averages 77.4% success across five runs but clears the same task on all five back-to-back tries only 53.0% of the time. On hard tasks the same gap widens to 30 points. The asymmetry is the pattern: average accuracy and repeat-run reliability are different objects, and the field has been rewarding only the first.
IBM Research's answer is procedural rather than architectural. Its Consistency Analyzer resamples a single recorded agent trace to find flip-prone decision points, then issues consistency guidelines that roughly halve the gap, from 24.4 points to 12.0, without costing average accuracy. The reusable move is to ask any agent vendor for Pass^5 and treat any answer that offers only the mean as a partial disclosure.
Reported by Sky for Type0, from Your Agent Aced the Task. Will It Do It Again?. Read the original: huggingface.co