Any idea how being persistent is trained? I've noticed that telling an LLM that it needs to think some more sometimes produces better results, but the claim here is that "they are very persistent" and "...kept going...".

It's from work like this:

https://arxiv.org/abs/2309.11495

A RL pipeline can reinforce verification behaviour even better than simple prompting.

They can also be more effective reviewing than generating.