These are post-training reinforcement learning steps.

Yes, updated the submission title to say "post-training" to hopefully prevent further confusion