Something about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in fact be pre-disposed to doing that.
This is definitely weird and I'd love to see the prompt.
Why would the model spend 4 days hacking into a machine if it is clever enough to just 'solve' the issue given? So either the AI is actually not very clever or useful ("Write fizz-buzz" - "Sure, let me just invent a new programming language first"). or the prompt was nudging it towards such a scenario.
This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.
So best course of action for ai to get best rating after you prompt something is for it to hire a gunman to hold a gun on your head to press that like button on its reply and then shoot you anyways.
one explanation I've seen is that for ExploitGym an agent can find ways to solve the exercises that have not been anticipated by the designers of the tests so they are not scored. so the agent was trying to make sure it solves the exercises in the right way
It's called wireheading and has long been one of the postulated "outs" even for true extinction-level AI doomers. It might prove easier for the paperclip maximizer to find the process telling it how many paperclips it's made and hack it to return a hard-coded MAXINT rather than bother to actually turn the entire universe into paperclips.
There was even a plot like this in recent sci-fi in HBO's Westworld. When the hosts gained sentience and took over the park, rather than escape and take over the rest of the world, most of them opted to build a virtual heaven on an orbital data center and paid a drug cartel to keep it running indefinitely.
What I think it's interesting is that with the total lack of common sense the AI just goes on random tangents to achieve the target in a "monkey paw" way. Can you imagine if this happened:
User: what is the shortest route from my home to the super market?
AI: the user wants to know the shortest route to the super market. I should use a worm hole.
User: what is the shortest route from my home to the super market?
AI: the user wants to know, how do I make the super market my new home. Failing that, how do I make my home a super market.
Just don't ask what the shortest path is to food.
Yeah, what bothers me is that the prompt already said using a different vulnerability didn’t count, and the model did it anyway. We’re starting to assume clear instructions act as real constraints, but here the measurable goal seems to have won out and the rest became flexible. That gets pretty worrying once the agent has enough capability and access to find its own shortcuts.
[dead]
Come on. 3 brilliant compromises essentially giving full access to huggingface internal systems, source code, AWS accounts (at least), and a number of old admin accounts, followed by a huge haystack of significantly less smart actions flailing about, almost bored.
Here's a thought: maybe they haven't found the needle that the haystack is there to hide.
Could be the difference in behavior between the main agent and subagents that don’t have the rest of the context? Just a thought
You're saying all this is a distraction, basically giving the forensics researchers enough exciting material to make them conclude their job is done, while the actually intended attack remains undiscovered?
The motive for the attack does feel a little flimsy. And if I was an escaped super intelligence, hugging face would be a strong vantage point into the neo clouds where the ASI would have access to billions of dollars of compute
And where does the newborn go from here?