That should be the easiest part of the loop. Frontier models are good judges of human preference and taste and should be excellent at prioritizing proposed questions. The generation part is much harder.