Independent agent research

Benchmarks that show their work.

We test frontier coding agents on real engineering tasks—and publish the methods, costs, failures, and evidence behind every result.

Current campaign

Controlled security campaign v1In progress and paused at a clean 11-block boundary: 22 terminal assignments, 50 pending, quality ungraded.

Our standard

01

Real work, not toy prompts

Tasks come from production repositories with authentic constraints, ambiguity, and failure modes.

02

Controls before conclusions

We preregister settings, isolate runs, preserve artifacts, and separate pilot evidence from controlled claims.

03

Every tradeoff visible

Quality sits beside time, tokens, cost, tool use, and variance. A single score never tells the whole story.