The BMS journal
Notes from the test bench.
Essays on evaluation design, agent economics, reproducibility, and what we learn while running the benchmarks.
MethodologyField note
What makes an agent benchmark worth trusting?
The controls, disclosures, and uncomfortable edge cases we think every serious coding-agent comparison should publish.
Cost is a capability metric, not a footnote
Why tokens, turns, and elapsed time belong next to quality scores—and what each one reveals about an agent.