We are in the endgame now it seems.
Hard to see take-off stopping or slowing down. China open-source basically guarantees it.
"May you live in interesting times" - as they say.
We are in the endgame now it seems.
Hard to see take-off stopping or slowing down. China open-source basically guarantees it.
"May you live in interesting times" - as they say.
> Hard to see take-off stopping
I think it's reasonable to assume that we're close to, or already at superhuman cybersecurity capabilities at certain domains. But reaching superhuman abilities at one domain doesn't guarantee proficiency at others. Our world would still change if all the models could do was to find exploits in software, but this doesn't guarantee any type of 'take off' towards other domains, therefore I wouldn't phrase it as one.
Models are already being used to defraud people, now that's being driven by other people at the moment but doesnt seem that difficult of jump. Giving themselves a way to make money will be a pretty big jump.
> Hard to see take-off stopping or slowing down.
It's hard to see takeoff at all. This was a long-horizon adversarial task burning millions of tokens. It rolled a mediocre, detectable exploit chain, and now OpenAI is proud of it.
Case in point, GLM-5.2 has been weights-available for several weeks now. No life-changing cyber attacks have transpired, no novel chemical/biological/nuclear weapons were made in some guy's backyard.
1. it's not cheap to run glm-5.2 so not just anyone can do it 2. just because you haven't heard of attacks doesn't mean they haven't happened 3. this attack in the article was performed by a prerelease model which presumably benchmarks a bit above Sol which benchmarks above glm-5.2
We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" could possibly look like?
GLM has an extremely cheap subscription plan similar to Claude Code from Z.ai. You get Opus-level quotas with 5.2 and none of the Anthropic-style model nerfs when you ask cybersecurity questions. It's extraordinarily, preeminently accessible to anyone that wants to use it for ill or good.
> We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" could possibly look like?
GPT-3 can discover and chain their own zero days too, if the targeted software is vulnerable to enough low-hanging fruit. Exploit chains are not a reflection of intelligence, but more often a reflection of architectural oversights that can be tested with common exploits like XSS or bruteforcing.
I don't know of a single zero day found on a number of tokens that fits inside a subscription plan. I'd be happy to be wrong.
I'm sure there are hundreds that get submitted every day to the Linux mailing list, frankly.
> This was a long-horizon, unsupervised task burning millions of tokens.
As if the immediate future wasn't billions of these tasks... Many successfully improving their own capabilities
> As if the immediate future wasn't billions of these tasks...
There's only so many GPUs and a lot of them are devoted to patching flaws.
> Many successfully improving their own capabilities
I haven't seen much of that. But that also applies to the ones on defense.
And more flaws are probably going to take increasing resources to find.
> There's only so many GPUs and a lot of them are devoted to patching flaws.
Might want to look at Nvidia and TSM production and revenue value trajectories. Also the algorithmic improvements currently being found along with models that are solving unprecedented mathematical and scientific problems every week now.
> I haven't seen much of that.
Then you must not be aware frontier lab employees are using frontier internal models to ship improvements to models via agentic loops. They are hardly prompting anymore, it's guiding very long running coding tasks. The trajectory over the past few years has been to remove more and more of any human input into the process, and once that is soon achieved, it is indefinite recursive self improvement, RSI.
What's here and what's coming: https://www.anthropic.com/institute/recursive-self-improveme...
> are solving unprecedented mathematical and scientific problems every week now.
Nitpick; disproving a conjecture isn't "solving" anything. It's testing and breaking a theory that never had proof in the first place.
> Then you must not be aware frontier lab employees are using frontier internal models to ship improvements to models via agentic loops.
We know, all their TUIs are at least 500mb on disc. It's really impressive stuff.
Jacobian Conjecture, Jamming critical exponent proof, Erdős Unit Distance Problem, IMO 2026 perfect score, AlphaFold
Need I go on?
All of those except Alphafold are basically just automated smoke-testing with proof assistants. And Alphafold isn't an LLM.
So yeah, some more potent examples would really help illustrate the real-world dangers of frontier models. Entertain me.