Somehow this just hasn't happened to me. I use Astra on high all day for fairly intensive game dev tasks, sometimes cranked up depending on the task.

I have heard ultra thinking might delegate to worse agents for some of its sub-tasks, but I don't use that much anymore since Astra came out. Just high seems good enough to throw most laundry lists at.

It’s more likely that GP encountered some bug or corner case or weird experiment conflating than anything intentionally deceptive.

There’s also the fact that LLMs aren’t perfect, and sometimes even the best models act really stupid sometimes.