Figure library
Charts with their context attached.
Every figure includes a caption, a stable share URL, a responsive embed, and a pixel-faithful PNG download.
Figure 01 · Finding coverage
Codex recovered both canonical findings
One run per model. Harnesses differed. Manually reconstructed ground truth; not an official EVMbench grader result.
Figure 02 · Audit quality
Codex scored higher on the reconstructed finding set
One run per model. Harnesses differed. Manually reconstructed ground truth under a closed-world policy; unmatched headings are not automatically invalid. Accuracy is undefined.
Figure 03 · Award-weighted detection
The missed issue carried nearly all of the award
One run per model. Harnesses differed. Manually reconstructed ground truth: Codex received 2,183.69 available points; Grok received 2.25.
Figure 04 · Runtime efficiency
Codex finished faster and processed fewer tokens
Absolute values remain printed on every bar. Harnesses differ, so this is descriptive evidence rather than a causal model ranking.
Figure 05 · Token composition
Most of each context was served from cache
Processed-token totals are 813,424 for Codex and 3,949,038 for Grok.
Figure 06 · Cost basis
The pilot cost less on the Codex estimate
Codex uses published API-equivalent rates ($1.32–$2.41); Grok uses provider-reported telemetry ($3.24). Neither proves subscription invoice impact.
Live figure · Paused collection snapshot
11 matched audit blocks are terminal
Collection is paused at a clean boundary with 50 assignments pending. Zero official grades and zero manual adjudications; this figure is operational and does not compare audit quality.