It turns out that LLMs, especially local LLMs, tend to hallucinate a lot when thinking about anything that's overtly fiddly or technical. This is even more the case when they're in a domain that isn't a natural part of their training data. If you have to "jailbreak" the model to get it to talk, you're so wildly out of the expected distribution that you'd be crazy to trust anything it says. It's basically making up stuff as it goes along. These are foundational issues with how the models are created, not something that a bad actor can just hack around.
(The biggest real safety issue in this kind of space is actually that the model might actively goad some unsuspecting victim into doing something incredibly dumb and dangerous to themselves as much as possibly others.
IIRC, there were reports of something vaguely similar happening IRL but involving casual mischief, not any kind of extreme attacks. And because nobody else seems to have managed to elicit the same actively goading verbiage from the model, it's implicitly suspected that the person involved was the one who introduced the problematic scenarios to begin with.)
This take belongs in 2024. It has been falsified multiple times but it never seems to go away.
Are you thinking about model capabilities in coding and math starting late 2025 or so? Those were intentionally boosted via automated RLVR, and there's nothing even loosely comparable to that in applied biology work, let alone in the speculative "helping a bad actor do something crazy" domain that the AI safety folks are worried about. You can't extrapolate from one to the other.
The implied concerns from sensible safety advocates are also about someone jailbreaking the latest proprietary AI frontier model for something like this (which is why their current guardrails are so extreme), not about toy local models.
I actually agree with you. Verifiable domains will have better performance.
But there’s nothing specially bad about LLMs that don’t allow it to work outside of its training set. It’s just that biology has to verify itself in physical realm and it’s a bit slower.
So yeah, I also don’t think some bad actor will find the secret to manufacturing a bio weapon using LLMs. But maybe these people think it’s possible. I’m skeptical but I’m going to also listen to the people who know it best.
> But maybe these people think it’s possible. I’m skeptical but I’m going to also listen to the people who know it best.
The problem is that in order for the scaremongering to make any kind of sense and for "stop frontier AI immediately" to be the right response (which is what the "AI safety" folks seem to be pushing for), you don't just need this to be possible in the abstract at some undetermined point in the future. You also need to argue that it will not be helpful for white-hat biosafety researchers (there will hopefully be several orders of magnitude more white-hat biosafety folks than attackers, with orders of magnitude more resources available) to red-team that exact scenario several months or even years in advance using their trusted access to unreleased super-smart AI, and thereby devise appropriate defenses with that same AI's help. That, if anything, is the most implausible part about this entire scenario.
>You also need to argue that it will not be helpful for white-hat biosafety researchers (there will hopefully be several orders of magnitude more white-hat biosafety folks than attackers, with orders of magnitude more resources available) to red-team that exact scenario several months or even years in advance
Sigh.
Attackers only need to win once. Defense needs to work every time.
A single wide scale attack affecting around 100k people or more will have your neighbors stomping on your face telling you to shut up, and to lock this shit down.
It's insane how you can watch a technology get better and better and better and come up idea that everything will remain the same. We are currently in the middle of development of the most powerful weapons on earth and you don't want to think about it because it's uncomfortable.