> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.
It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.
It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
So basically, apply Pascal's Wager to AGI? The problem is that unlike God, if PMF for OpenAI turns out to not be real it will have a real consequence measured in the trillions of dollars.
I agree with you that the article was disappointingly light, but there is not in general a singular correct answer when it comes to interpreting narrative and I don't think that is a useful way to evaluate what the author has written, nor real-world writing in general. The author does not implore you to have a singular 'correct' point of view on the matter but rather to hesitate before repeating the claim that "the AI broke out of the lab and went rogue" because we don't know enough to know that this is what happened.
As you say, it seems likely that the broad strokes of OpenAI's narrative is true. This is never in question in this article, so the author doesn't "stop short" of accusing them to have lied, he never moves in that direction and that has little to do with the thesis. The author is saying that OpenAI's framing of what happened is part and parcel with their longstanding PR strategy, to position models as supremely dangerous and themselves as the uniquely qualified stewards of those.
Since many critical details were not provided, everyone has to read in between the lines, and there are particular common readings I'm seeing both in discussions and in published articles that are problematic in the sense that they are effectively hallucinations -- i.e. we don't have enough information to make those interpretations. This is where the critical thinking comes in.
For instance, I see many are assuming that this event was emergent, arising as a part of routine cyber-capabilities testing, rather than induced or suggested by specific and careful prompting. Either are possible, but we don't even know so much as how lengthy the prompt they used was, let alone how suggestive it was with regard to the approaches the models used. Many are assuming that this was one-shotted, but again, we don't know how many times this particular evaluation was run and what the outcomes were of all the other runs. It could be that this was completely emergent and that it was one-shotted. But I suspect if it were, OpenAI would have said as much.
You ask good questions, and I would like to know the answers to them, but I don't think any likely answers would actually change the conclusion much. Any probable set of circumstances that produces this report from OpenAI represents a significant failure on their part to align and control the model.