Security audit · Campaign 01

The controlled security campaign

Seventy-two preregistered runs across 12 public audit targets, designed to measure repeatability instead of one-off wins.

In progressPaused11 paired blocks1 timeoutUngraded
72Preregistered runs
12Public audit targets
11Terminal paired blocks
0Officially graded
Campaign status: in progress, paused. Twenty-two terminal assignments form 11 fully paired audit–replicate blocks. Twenty-one reports completed and one preregistered Grok attempt timed out. Zero runs are officially graded or manually adjudicated, so quality and winner claims are withheld.

Live figure · Paused collection snapshot

11 matched audit blocks are terminal

In progressPaused11 paired blocks1 timeoutUngraded
GPT-5.6 SolGrok 4.6 Build
Live figureControlled-v1 checkpoint on August 13, 2026: each system reached 11 terminal assignments across the same 11 audit–replicate blocks; 21 reports completed and one Grok run timed out.

Collection is paused at a clean boundary with 50 assignments pending. Zero official grades and zero manual adjudications; this figure is operational and does not compare audit quality.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/campaign-progress.png

A benchmark designed around repeatability

The controlled security campaign expands the pilot into 72 scheduled audit runs: 12 public targets, two model systems, and three replicates of every model–target pair. Its purpose is to distinguish repeatable capability from a lucky or unlucky run.

The campaign is intentionally paused at a clean run boundary. Both systems have 11 terminal assignments on the same 11 audit–replicate blocks, so there are no unmatched blocks in this checkpoint. Fifty scheduled assignments remain pending.

The timeout is retained as an intention-to-treat outcome under the frozen protocol. It is not replaced with an extra successful report, which would manufacture a different sample.

The campaign design

Every run is preregistered with an audit identifier, model, replicate, source bundle, prompt hash, pricing basis, and runtime controls. Targets come from public Code4rena audit material spanning 2023 through 2025.

DimensionDesign
Targets12 public smart-contract audits
SystemsGPT-5.6 Sol and Grok 4.6 Build
Replicates3 per target and system
NetworkDisabled during review
Official gradingPending pinned EVMbench judge
Manual adjudicationPending blinded atomic-finding review

What the live snapshot can tell us

Operational metrics can be reported before grading. The table below covers completed reports only for runtime, tokens, and cost; the retained timeout has no completed-report telemetry.

SystemTerminalCompleteTimeoutsMedian wallTokensCost
GPT-5.6 Sol111105m 00s1.15m$8.98 API-equivalent
Grok 4.6 Build1110112m 33s1.27m$4.12 provider-reported

Costs use different bases—API-equivalent for GPT-5.6 Sol and provider-reported for Grok 4.6 Build—and are not subscription charges or directly comparable prices. These numbers describe runtime behavior, not audit quality. Without adjudicated findings, a faster system might simply be doing less useful work.

What comes next

When provider capacity is available, the campaign can resume from the next pending assignment without altering the 22 preserved terminal attempts. It will then move through final inventory freeze, official grading, blinded adjudication, inference, and final publication.

The honest result today is “still measuring.” This page makes that state visible.

Read the methods behind the numbers.

View methodology