Our goal is not to crown a universal winner. It is to produce useful, falsifiable evidence about how coding agents behave under a specific set of constraints.
Principles
Use real tasks. We prefer production repositories, public audit targets, and work that has an externally meaningful definition of success. Toy tasks are easier to grade, but they erase the context management and judgment that make agents useful—or dangerous—in practice.
Control what can be controlled. Models receive equivalent workspaces, prompts, tool access, time boundaries, and network policy. We record observed model identifiers and runtime versions rather than trusting labels alone.
We treat every benchmark result as a statement with a scope, not as a leaderboard truth.
Protocol
| Stage | What is fixed | What is recorded |
|---|---|---|
| Prepare | Commit, scope, prompt, controls | Hashes and environment |
| Run | Fresh workspace, tool policy | Time, tokens, tool events |
| Grade | Rubric and ground truth | Matches, misses, disputes |
| Publish | Disclosure schema | Results and limitations |
Grading
Security findings are matched to canonical issues by root cause and impact, not by title similarity. Precision, recall, F1, and severity agreement are reported separately. Where no complete true-negative universe exists, we leave “accuracy” undefined rather than manufacturing one.
Reporting
Every publication separates provider-reported cost from API-equivalent estimates and subscription charges. It also puts latency and resource use beside outcome quality. Draft campaigns are visibly marked in progress until adjudication is complete.
Limitations
Agent systems change quickly. A result describes the tested model build, harness, date, and task set; it does not guarantee future behavior. Small samples are pilots. Controlled campaigns use replicates specifically because agent outcomes are stochastic.