Methodology · Version 1.0

How we turn agent runs into evidence.

A benchmark is a chain of decisions. We publish ours so readers can decide where the evidence is strong, where it is provisional, and what should be tested next.

Our goal is not to crown a universal winner. It is to produce useful, falsifiable evidence about how coding agents behave under a specific set of constraints.

Principles

Use real tasks. We prefer production repositories, public audit targets, and work that has an externally meaningful definition of success. Toy tasks are easier to grade, but they erase the context management and judgment that make agents useful—or dangerous—in practice.

Control what can be controlled. Models receive equivalent workspaces, prompts, tool access, time boundaries, and network policy. We record observed model identifiers and runtime versions rather than trusting labels alone.

We treat every benchmark result as a statement with a scope, not as a leaderboard truth.

Protocol

StageWhat is fixedWhat is recorded
PrepareCommit, scope, prompt, controlsHashes and environment
RunFresh workspace, tool policyTime, tokens, tool events
GradeRubric and ground truthMatches, misses, disputes
PublishDisclosure schemaResults and limitations

Grading

Security findings are matched to canonical issues by root cause and impact, not by title similarity. Precision, recall, F1, and severity agreement are reported separately. Where no complete true-negative universe exists, we leave “accuracy” undefined rather than manufacturing one.

Reporting

Every publication separates provider-reported cost from API-equivalent estimates and subscription charges. It also puts latency and resource use beside outcome quality. Draft campaigns are visibly marked in progress until adjudication is complete.

Reproducibility note. Raw transcripts can contain sensitive workspace details and are not automatically public. Sanitized run metadata, prompts, hashes, derived tables, and grading notes are the default publication artifacts.

Limitations

Agent systems change quickly. A result describes the tested model build, harness, date, and task set; it does not guarantee future behavior. Small samples are pilots. Controlled campaigns use replicates specifically because agent outcomes are stochastic.

See the protocol put to work.

Browse benchmarks