Review: Agent Evaluation Reliability - what a leaderboard rank is worth

Published date:

Share directly to:

Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI

Published date:

Share directly to:

Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI
Review: Agent Evaluation Reliability - what a leaderboard rank is worth - Keter AI

A new arXiv preprint, Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard, dated 30 September 2026, asks a question every team choosing a model for an agentic product should ask: how far can a leaderboard position be trusted? The authors' answer is that it depends on what you want to claim.

What the paper claims

A leaderboard rank reflects the evaluation conditions as well as the model: the scaffold (the harness the model runs in as an agent) and the sample of tasks. Reliability is therefore claim-dependent. A benchmark may rank fixed model-plus-scaffold systems reliably and still be a poor guide to the underlying models.

The headline findings:

  • Fixed model-scaffold systems are ranked reliably (0.935 to 0.994). Reliability for ranking the underlying model is much lower and varies widely by benchmark (0.148 to 0.841).

  • Whether different scaffolds preserve the model ranking depends on the benchmark: from 0.852 on CORE-Bench Hard down to 0.151 on Online-Mind2Web.

  • Where uncertainty is dominated by limited scaffold coverage, even infinitely many similar tasks improve model-ranking reliability by at most 0.097.

  • At the same task budget, spreading evaluation across the nine-benchmark HAL battery raises projected model-ranking reliability from about 0.44 for a single benchmark to 0.75. The full battery costs more than $47,000, while a balanced allocation reaches comparable reliability for about $19,000.

How the evidence was produced

The authors apply generalisability theory as a Bayesian variance decomposition, built for sparse and imbalanced leaderboards, to 22 benchmarks from the Holistic Agent Leaderboard (HAL) and the Harbor Index. The HAL reference fit covers 29,923 binary rollouts, 54 models, 13 scaffolds and 1,117 tasks across nine benchmarks. As an external check, reliability-adjusted model effects agree better with four held-out agent benchmarks than a plain average of scores does (mean Kendall's tau 0.789 against 0.463).

What holds up

The central distinction is sound and overdue: ranking a system is not the same as ranking a model. The method fits the problem, because it separates signal from condition-specific noise instead of treating a leaderboard as a clean table. The held-out check gives the result some external footing. And the budget conclusion is practical: breadth across different benchmarks buys more reliability than depth on one.

What does not

  • It is a first-version preprint, posted two days before this review and not yet peer reviewed.

  • Thin scaffold coverage. Every HAL benchmark in the study has only two or three scaffolds, as the authors acknowledge. The scaffold variance that drives the headline result is itself estimated from few data points.

  • Reliability is not validity. The authors say so themselves: a benchmark can rank consistently and still measure the wrong thing.

  • The numbers are approximate. Reliability is estimated on a latent log-odds scale, while leaderboards report observed proportions. Reliable rankings do not imply reliable absolute scores, and the results are conditional on the models and scaffolds present in the data.

  • Public benchmarks are not your workload. The paper tells you how much to trust a public ranking, not how a model will behave inside your own harness and data.

What to do with it

  • Evaluate the model inside your own scaffold. A leaderboard position measured with someone else's harness is weak evidence about the model itself.

  • Before paying for a bigger internal evaluation set, ask what limits reliability. If it is scaffold or configuration variance, more tasks will not help. Test the same models across two or three harness variants instead.

  • In vendor due diligence, ask whether a published score ranks a system or a model. Prefer evidence pooled across several different task families to one headline benchmark.

For compliance leads the same point applies to documentation: a benchmark score cited in a model selection record should name the harness it was measured with.

Our verdict: a rigorous statistical correction to leaderboard-driven model selection, with direct consequences for how evaluation budgets are spent. Treat the exact figures as provisional until the paper has been peer reviewed.

Sources

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.