Great idea, we’d love to run those benchmarks! Any in particular you trust?
Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
Great idea, we’d love to run those benchmarks! Any in particular you trust?
Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
I’m not really sure, I can’t say I trust any kind of synthetic AI benchmarks.
> Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
Permissions issues would be the most common - models collaborating with eachother on a task seem to try to convince eachother they have either more or less permission to do things than they actually do. Especially when transitioning between creating a plan and executing the plan.
Another is deciding that there is a limit to the “loops” they are allowed to run to iterate on something. In many cases I have set an explicit goal, and come back to an agent stopped and reporting that it has hit the “maximum allowable loops of [insert arbitrary number that changes every time].”
Now that I think about it further, I believe the other examples I have also all fall into the models hallucinating the presence of control instructions, or attempting repeatedly to violate permission boundaries that a single model’s harness instructions would usually guide it away from re-attempting.
> I’m not really sure, I can’t say I trust any kind of synthetic AI benchmarks.
Fair enough but what exactly were you thinking when you said:
> I would like to see this (and any model router, frankly) benchmarked against two things
And thanks for sharing about failure cases! Those do sound like strange harness-level things, honestly we haven't seen failure cases like that crop up in our own usage & testing.