Figure 02 · Audit quality

Codex scored higher on the reconstructed finding set

GPT-5.6 SolGrok 4.6 Build
Figure 02Recall, manual precision, and F1 for submitted audit findings. Higher is better.

One run per model. Harnesses differed. Manually reconstructed ground truth under a closed-world policy; unmatched headings are not automatically invalid. Accuracy is undefined.

BMSbms.kyokasuigetsu.xyz
View source at BMS Static image: https://bms.kyokasuigetsu.xyz/charts/quality-metrics.png