I was excited until I saw cost was only provided as the median. Your provider will bill you for all your tasks and one can get back to that total from the mean by multiplying by the number of tasks. This isn't possible with the median and I suspect the median is likely below the mean so this understates the actual costs.
We initially tried using the mean (average), but a few extreme outliers caused Claude Code's cost to look far higher than it typically is.
Claude Code may simply be best used with Anthropic's models and quite bad with Kimi. An alternate solution is to remove Claude Code from the diagram if its so far off from the others that it causes scaling problems.
I was looking at the results JSON and it looks like there is only one run of each task with each harness. Since these are disparate tasks using the median means that the headline cost is the cost of one specific task for each harness (or the mean of two it looks like in the case of Exo Harness; I couldn't spot the one task that lined up with the headline cost), but the same task isn't used for the headline cost for each harness. It's not the same as picking a task at random to use as the headline task, but its in the ballpark.
[flagged]