Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)
Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)
They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.
the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.
It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?