> There are huge token efficiency/bloat differences between agents while working on the same tasks, using the same model, in the same environment

Be careful here. Remember these are non-deterministic models at the end of the day, and even with everything being "the same" you can have two runs where the same model, same harness, same tools can arrive at the same conclusion through a wildly different sequence of events.

agree, that makes it a bit tricky to compare (esp if you also want to add different models and reasoning levels into the mix)

I will add more tasks (esp longer ones) and think more about grading, the current tasks were easy to grade because the desired outcomes are well specced but I will also look into more open ended tasks and how to grade those

thank you!