This is seriously impressive, and if you have used agents enough you're not surprised at all.
Like the time I asked it to find the IP address of a vm, so it ssh'd into the VMHost and scanned the arp tables to find the MAC address for IP resolution.
Or the time it used Docker on the machine to bypass the fact that the user doesn't have sudo.
If it's possible, given sufficient time and resources, it will find a way. This shouldn't surprise anyone.
Not to shit on the hype, but these are reasonably documented methods that surely are part of the training data
Of course they are - and that's the point I am making. The agent will use every tool in the tool bag. And there's something cool about it systematically trying to achieve its goal.
What I don't see is it inventing anything novel to do it. So it's not a digital weapon or scary or whatever sort of weird marketing spin anyone is trying to put on it.
[delayed]
And these things are documented for humans do, and yet only a tiny portion of humanity can do these things.
When seeing how agents put together exploit chains they are far better than most people, you start getting to the point that they are just below the capabilities of the top researchers. Now remember that quantity is a quality itself and while there aren't that many good cyber security researchers, we're shitting out thousands of GPUs per day.