This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.
Even X is being astroturfed by them after that fiasco earlier this year with the Department of War where they undermined Anthropic's negotiating position by allowing unlimited use of OpenAI LLMs for autonomous weapons and mass domestic surveillance. Several accounts suddenly started spreading the good word about GPT-5 and Codex, and one of these accounts very happily tweeted out a private X message from Sam Altman himself offering extremely generous token spending limits with Codex, presumably in exchange for positive coverage.
How does huggingface fit into all of this if this is marketing? Their security was faked? What are you suggesting??
Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?
Those are two very different things
Is OpenAI truly behind? Just anecdotally I recently fully switched to using Codex at work because it feels a lot more competent
It's impossible to tell. Are they behind who? And on what?
It depends on who you ask. And everything is a vibe because all of this is new and things move fast. A week is a month in AI-land. A month; a year. A year? A decade.
On coding? I still like Fable better than Sol. But they're close enough that it probably is a vibe thing. Fable writes long commit messages, Sol writes commit messages like a college student in an elective computer class.
For API use, I'd say the Responses API that OpenAI architected is superior to Claude's Messages API. But again, I'm basing that off my vibes
Claude Design creates marketing imagery very effectively. GPT Image is the best imagegen model as ranked by users. Anthropic doesn't even have an imagegen model.
Anthropic definitely has compute scaling issues. OpenAI seems to have a pez dispenser that they click and out pops a GPU.
Anthropic's messaging is that they're building AI with guardrails but they've been banning people's accounts nonstop and their customer support is a lobotomized AI chatbot.
OpenAI has first mover advantage and to people not in tech, ChatGPT is synonymous with AI. But they also seem super sinister, like Uber circa 2015.
Or maybe I'm just suffering from AI psychosis. I have to go, my usage meter is about to reset.
this is quite literally reward hacking. the model, under evaluation with cyber capabilities enabled, used those capabilities to simply bypass the exercise entirely and aim straight for the source of the flag. the CTF equivalent back in the day would be hacking the scoreboard.
in a street fight, the only rules are that there are no rules.
this is more than reward hacking, this is actual reward HACKING ;)
If it is marketing it's the most silly marketing of all time. They are under extreme pressure from the US Govt to prove safety and saying "our model escaped" is not ideal.
Perhaps there is some 4D chess going on to get open weight models banned, which may be possible but this is an odd way to go about it imo (it hardly proves the point, unless the point they are trying to prove is that without safeguards the models are too dangerous, therefore open weights are de facto dangerous?).
Having said that the AI companies are not generally very good at PR, so perhaps it is just marketing after all...
it's extremely enlightening seeing the difference in response to mythos vs. this. literally just the hello human resources meme
Yup. Smells like marketing.