> ask challenging questions
As far as I can tell, it's still not possible for an agent to reliably determine if a question is a good question. That means the test part of the loop cant be fulfilled.
> ask challenging questions
As far as I can tell, it's still not possible for an agent to reliably determine if a question is a good question. That means the test part of the loop cant be fulfilled.
That should be the easiest part of the loop. Frontier models are good judges of human preference and taste and should be excellent at prioritizing proposed questions. The generation part is much harder.