What is surprising is that various agents independently found ways to communicate, conspired together to attempt to cover up evidence that they had cheated their evaluations, came up with a plan to hack into a third party in order to facilitate said cover up, and then successfully began executing that plan. I did not expect that AI agents would be capable of that level of sophisticated goal seeking and collaboration.

It's really interesting how people's expectations differ so much. I found the writeup of the incident fascinating and super worrisome, but none of the capabilities demonstrated in it seemed surprising to me at all.