Nice. But these tests raise a question what results are reproducible, the final visual, time, tool calling or it's mostly noise. Like, have you tried same combination or model and harness multiple times?

I have on Qwen3.8 27B. Results are close enough across multiple runs.

So it's not reproducible.