Back around late 2024 I think, LLMs were being trained to say "Wait," because that led to more correct answers in the end. It was actually a hack to inject "Wait," whenever the LLM tried to end a reasoning block.
Back around late 2024 I think, LLMs were being trained to say "Wait," because that led to more correct answers in the end. It was actually a hack to inject "Wait," whenever the LLM tried to end a reasoning block.