The training here is RL training, the rollouts there are not different from inference and have access to the same tools as regular inference.
The training here is RL training, the rollouts there are not different from inference and have access to the same tools as regular inference.