Shareable benchmark figure
Codex recovered both canonical findings
Canonical findings matched in the PoolTogether V5 audit pilot. Codex matched both reconstructed ground-truth issues; Grok matched one.
Figure 01 · Finding coverage
Codex recovered both canonical findings
GPT-5.6 SolGrok 4.6 Build
One run per model. Harnesses differed. Manually reconstructed ground truth; not an official EVMbench grader result.