The benchmark index
Real tasks. Full receipts.
Controlled evaluations of frontier coding agents on production engineering work, with methods, artifacts, and limitations published beside the result.
01
Codex vs. Grok on a real smart-contract audit
A one-target case study of frontier coding agents on PoolTogether V5, measuring detection, precision, speed, tokens, and cost across differing harnesses.
GPT-5.6 SolGrok 4.6 Build
02
The controlled security campaign
Seventy-two preregistered runs across 12 public audit targets, designed to measure repeatability instead of one-off wins.
GPT-5.6 SolGrok 4.6 Build