Shareable benchmark figure

The missed issue carried nearly all of the award

Award-weighted detection score against the reconstructed canonical findings. Higher is better.

Figure 03 · Award-weighted detection

The missed issue carried nearly all of the award

GPT-5.6 SolGrok 4.6 Build
Figure 03Award-weighted detection score against the reconstructed canonical findings. Higher is better.

One run per model. Harnesses differed. Manually reconstructed ground truth: Codex received 2,183.69 available points; Grok received 2.25.

BMSbms.kyokasuigetsu.xyz
Share figure
Static image: https://bms.kyokasuigetsu.xyz/charts/detect-score.png