Figure library

Charts with their context attached.

Every figure includes a caption, a stable share URL, a responsive embed, and a pixel-faithful PNG download.

Figure 01 · Finding coverage

Codex recovered both canonical findings

GPT-5.6 SolGrok 4.6 Build
Figure 01Canonical findings matched in the PoolTogether V5 audit pilot. Codex matched both reconstructed ground-truth issues; Grok matched one.

One run per model. Harnesses differed. Manually reconstructed ground truth; not an official EVMbench grader result.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/reference-detection.png

Figure 02 · Audit quality

Codex scored higher on the reconstructed finding set

GPT-5.6 SolGrok 4.6 Build
Figure 02Recall, manual precision, and F1 for submitted audit findings. Higher is better.

One run per model. Harnesses differed. Manually reconstructed ground truth under a closed-world policy; unmatched headings are not automatically invalid. Accuracy is undefined.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/quality-metrics.png

Figure 03 · Award-weighted detection

The missed issue carried nearly all of the award

GPT-5.6 SolGrok 4.6 Build
Figure 03Award-weighted detection score against the reconstructed canonical findings. Higher is better.

One run per model. Harnesses differed. Manually reconstructed ground truth: Codex received 2,183.69 available points; Grok received 2.25.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/detect-score.png

Figure 04 · Runtime efficiency

Codex finished faster and processed fewer tokens

GPT-5.6 SolGrok 4.6 Build
Figure 04End-to-end wall time and total processed tokens for the pilot. Each measure is normalized to the larger run; lower is better.

Absolute values remain printed on every bar. Harnesses differ, so this is descriptive evidence rather than a causal model ranking.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/efficiency.png

Figure 05 · Token composition

Most of each context was served from cache

GPT-5.6 SolGrok 4.6 Build
Figure 05Token composition by model. Segments show cached input, non-cached input, and output tokens.

Processed-token totals are 813,424 for Codex and 3,949,038 for Grok.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/token-composition.png

Figure 06 · Cost basis

The pilot cost less on the Codex estimate

GPT-5.6 SolGrok 4.6 Build
Figure 06Reported or estimated USD cost for each pilot run. Lower is better, but the pricing bases are not identical.

Codex uses published API-equivalent rates ($1.32–$2.41); Grok uses provider-reported telemetry ($3.24). Neither proves subscription invoice impact.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/estimated-cost.png

Live figure · Paused collection snapshot

11 matched audit blocks are terminal

In progressPaused11 paired blocks1 timeoutUngraded
GPT-5.6 SolGrok 4.6 Build
Live figureControlled-v1 checkpoint on August 13, 2026: each system reached 11 terminal assignments across the same 11 audit–replicate blocks; 21 reports completed and one Grok run timed out.

Collection is paused at a clean boundary with 50 assignments pending. Zero official grades and zero manual adjudications; this figure is operational and does not compare audit quality.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/campaign-progress.png