I've got an application for this. Will try and get some benchmarks in the next few weeks. We're definitely attempting to converge on a quality inference without wasting tokens/time.
For now, we're passing the first pass to the humans, and that's roughly 80 for 20, but to get the last 20 without blowing up our token spend? This might be the approach.
Thanks for the write up comment, btw.
Hey gavmor, thanks for your comment! Let me know if you need any assistance, I would be happy to help or get any feedback at dave@loopgain.ai