Why BMS exists
Agent demos are abundant. Comparable evidence is not. We built BMS to make serious evaluation readable: rigorous enough for builders, clear enough for anyone deciding which systems to trust.
The name is deliberately plain. BMS is a bench, not a stage—a place to put systems under load, measure what happens, and keep the failed parts on the table.
What we publish
Our benchmarks focus on consequential engineering work: security review, repository-scale implementation, debugging, and other tasks where correctness matters. The journal covers the messy work around those tests—protocol design, grading, agent economics, and surprises that do not fit in a chart.
Independence and disclosure
BMS is produced by Kyokasuigetsu. We do not accept model-vendor control over methods, grading, or conclusions. If a benchmark is sponsored or credits are provided, that relationship will be disclosed on the publication itself.