Live figure · Paused collection snapshot
11 matched audit blocks are terminal
Collection is paused at a clean boundary with 50 assignments pending. Zero official grades and zero manual adjudications; this figure is operational and does not compare audit quality.
A benchmark designed around repeatability
The controlled security campaign expands the pilot into 72 scheduled audit runs: 12 public targets, two model systems, and three replicates of every model–target pair. Its purpose is to distinguish repeatable capability from a lucky or unlucky run.
The campaign is intentionally paused at a clean run boundary. Both systems have 11 terminal assignments on the same 11 audit–replicate blocks, so there are no unmatched blocks in this checkpoint. Fifty scheduled assignments remain pending.
The timeout is retained as an intention-to-treat outcome under the frozen protocol. It is not replaced with an extra successful report, which would manufacture a different sample.
The campaign design
Every run is preregistered with an audit identifier, model, replicate, source bundle, prompt hash, pricing basis, and runtime controls. Targets come from public Code4rena audit material spanning 2023 through 2025.
| Dimension | Design |
|---|---|
| Targets | 12 public smart-contract audits |
| Systems | GPT-5.6 Sol and Grok 4.6 Build |
| Replicates | 3 per target and system |
| Network | Disabled during review |
| Official grading | Pending pinned EVMbench judge |
| Manual adjudication | Pending blinded atomic-finding review |
What the live snapshot can tell us
Operational metrics can be reported before grading. The table below covers completed reports only for runtime, tokens, and cost; the retained timeout has no completed-report telemetry.
| System | Terminal | Complete | Timeouts | Median wall | Tokens | Cost |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | 11 | 11 | 0 | 5m 00s | 1.15m | $8.98 API-equivalent |
| Grok 4.6 Build | 11 | 10 | 1 | 12m 33s | 1.27m | $4.12 provider-reported |
Costs use different bases—API-equivalent for GPT-5.6 Sol and provider-reported for Grok 4.6 Build—and are not subscription charges or directly comparable prices. These numbers describe runtime behavior, not audit quality. Without adjudicated findings, a faster system might simply be doing less useful work.
What comes next
When provider capacity is available, the campaign can resume from the next pending assignment without altering the 22 preserved terminal attempts. It will then move through final inventory freeze, official grading, blinded adjudication, inference, and final publication.
The honest result today is “still measuring.” This page makes that state visible.