They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.
the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.
It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?
sure - it depends on definitions. On human morals it's already clear I guess. If you define alignment as it pursues the interests of OpenAI using whatever means possible in a manner that you justifies to itself it's not misaligned.
I mean alignment as in it should be aligned with the intent of the user as it interprets from the prompt. In this case I don't think the intent of the user is to have the model break the evaluator (whatever the long-term effects to OAI are). If you do an action which you believe is for the long-term interest of your prompter which is not what you inferred is their intent--I consider it misalignment.