Methodology

What makes an agent benchmark worth trusting?

The controls, disclosures, and uncomfortable edge cases we think every serious coding-agent comparison should publish.

A benchmark earns trust when a skeptical reader can reconstruct the claim: what happened, under which conditions, how it was judged, and what evidence would change the conclusion.

That sounds obvious. In practice, agent comparisons often compress an entire experimental system into one percentage and a logo wall.

Start with the claim, not the leaderboard

“Model A scored 73” is not yet useful. On what task distribution? With which tools? At what budget? Scored by whom? A defensible benchmark begins with a narrow claim and keeps the scope visible.

Good benchmarks make it hard to accidentally believe more than the experiment showed.

Our PoolTogether pilot, for example, supports a claim about two observed agent runs on one smart-contract audit. It does not support a claim about every security repository or every future model build.

Publish controls as part of the result

Agent capability is partly harness capability. Network access, subagents, context windows, tool permissions, retry behavior, and time limits can all dominate the result. Hiding these details makes comparison impossible.

We publish the prompt hash, observed model identifier, reasoning setting, workspace policy, tool restrictions, runtime version, task commit, and cost basis. When harnesses differ, that difference belongs in the headline limitations.

Show uncertainty before it is flattering

Agents are stochastic systems. One run is a case study. Replicates let us see variance, task interaction, and failure rates. Until replicated evidence is available, we label the work a pilot and resist ranking language.

Incomplete campaigns remain visibly incomplete. Operational data can be published live, but outcome metrics wait for adjudication.

Keep the evidence trail

Reproducibility does not require dumping every raw transcript onto the public internet. It does require preserving inputs, hashes, structured outputs, grading decisions, and enough provenance to audit the publication.

Our rule of thumb: publish the smallest safe artifact set that lets another researcher challenge the result without exposing secrets or irrelevant workspace data.

Finally, corrections need a visible path. A benchmark should be a maintained research object, not a frozen marketing page.

See the ideas under test.

Explore benchmarks