Figure 03 · Award-weighted detection
The missed issue carried nearly all of the award
GPT-5.6 SolGrok 4.6 Build
One run per model. Harnesses differed. Manually reconstructed ground truth: Codex received 2,183.69 available points; Grok received 2.25.
Figure 03 · Award-weighted detection
One run per model. Harnesses differed. Manually reconstructed ground truth: Codex received 2,183.69 available points; Grok received 2.25.