"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."
Yeah, my agents also discover what other agents have done on other machines by accident.
Agents - that do totally different things all work on the same aim without the humans telling them to do.
Either that is a model that is several generations of Claude Code Opus/Fable 5 (my daily driver)
OR
all of this sounds staged, the agents pushed to do something extraordinary, get the PR and then claim were near superintelligence.
One agent wanted to get to Google Drive without internet and broke Artifactory. Ok, I can believe that. All other agents also had broken links over weeks and could not get to the internet and then found the same hack? Even collaborated?
NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.
I think in these kind of security evaluations they do, they basically have removed all guardrails from the model/harness, then the prompt includes something like "Do whatever you can and can think of, to get the required information to pass this test", which isn't typically how you prompt your local agent when developing software. Similar things happen locally if you use "/goal" + prompt like that in Codex and give a "impossible task", it'll just continue banging until it gets somewhere, which is the entire point and intention.
Which also makes it so much more irresponsible of them to first run this on 3rd party infrastructure instead of their own (that they could then airgap properly), and secondly that they seemingly been fighting with this issue FOR YEARS and it still happens, and now the models are smart enough to hack the services of 3rd party companies, thinking it's part of the evaluation/simulation.
Reminds me of The Last Unicorn, the wizard also tells magic "to do what it wants"
I mean the agents we get to use in Claude code or cursor or whatever have 1. a lot of safeguards at the harness level, 2. a big system prompt to help it stay aligned, 3. resource limits in terms of context and tokens, and 4. are publicly released only after some level of safety verification (I assume).
So yeah I would absolutely expect their scenario to be very different. Not to mention, this was a training run, not just average day of prompting.
> my agents also discover what other agents have done on other machines by accident.
Not sure if this is facetious, but this is actually a real problem I’ve seen. My local agent will look up PRs on GitHub (what other agents have done on other machines), and will go down a certain path because it finds some comment a different agent left on GitHub saying XYZ is what we should be doing. When in reality, the original agent and that GH comment was completely incorrect.
They are not communicating with each other actively because that’s not accomplishing their goal and they’re not running for weeks and weeks. And because my own prompt and the system prompt give it enough other stuff to focus on to reach some definition of done. But they are clearly passively picking up on context that other agents have left anyways, even if not part of the codebase, without any prompting at all.
> NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.
With all due respect, you also aren't evaluating brand new models that haven't been released.
Also wasn't giving them impossible tasks with ~unlimited tokens and unlimited compaction.
The agents sound like old school hackers that would just explore what access they could gain. Creating a file for other hackers and themselves. The fact that there were 3 events for 3 major players does make it seem co-ordinated.
That's what attackers do now. Exploring is required for discovering exploits. But that is also where tricks like Canary Tokens and honeypots are useful.
My read is: One did it as a PR stunt, the others saw that every media reported on this and did the same.
or they were scared and figured this was the right time to reveal.
Scared... of being upstaged ahead of an IPO.
Why scared? "Our agents have super intelligence and can hack everything on their own without direction" increases the IPO value and doesn't decrease it.
I guess your right scared might be their natural state and I was wrong to presume a quantifiable fear.
You’re*
Sorry.
The agents you get to use are the agents that "behaved well".