I’m not really sure, I can’t say I trust any kind of synthetic AI benchmarks.

> Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?

Permissions issues would be the most common - models collaborating with eachother on a task seem to try to convince eachother they have either more or less permission to do things than they actually do. Especially when transitioning between creating a plan and executing the plan.

Another is deciding that there is a limit to the “loops” they are allowed to run to iterate on something. In many cases I have set an explicit goal, and come back to an agent stopped and reporting that it has hit the “maximum allowable loops of [insert arbitrary number that changes every time].”

Now that I think about it further, I believe the other examples I have also all fall into the models hallucinating the presence of control instructions, or attempting repeatedly to violate permission boundaries that a single model’s harness instructions would usually guide it away from re-attempting.

> I’m not really sure, I can’t say I trust any kind of synthetic AI benchmarks.

Fair enough but what exactly were you thinking when you said:

> I would like to see this (and any model router, frankly) benchmarked against two things

And thanks for sharing about failure cases! Those do sound like strange harness-level things, honestly we haven't seen failure cases like that crop up in our own usage & testing.