Training run was reinforcement learning. It's at 10:10 in the video.

The speaker handwaves that one model found the RCE and then another model found a way to communicate via a message board.

Communication via a message board is sure to be in the training via e.g.some lesswrong scenario or similar or previous RL.

I don't find it really interesting because it is always "the agent found this and that". We don't know what has been RL'd before. We don't have the setup. We don't know if there was previous RL training on breakout scenarios.

It isn't science, more like a computer game.

Modern medicine evolved in much the same manner. Hand wavy practitioners copying each other without rigorous verification of efficacy that resulted in many lives lost.

That’s why medical research has so many hoops to jump through.