Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?
Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?
They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.
Precisely. "Aw jeez, we finally built the T-1000, but all it wants to do is kill John Connor – just like we warned! Why did I give it live ammunition and unsupervised time machine access?"
He wouldn't be the first reckless CEO...
They said: AI is becoming dangerously autonomous and capable. Proof of today's breach. Crowd "hey why didn't you say so, c'mon it's marketing". Them "we said so".
“Never attribute to malice that which is adequately explained by stupidity.” (or carelessness in this case)
I would attribute it to profit motive instead of either stupidity or malice.
FWIW, I used to love this phrase but over recent years have come to understand it is quite damaging. We live in a society where evil frequently hides behind a ‘stupid’ label, and people bring this quote up to defend or soften actions that are indeed done out of specific malicious intent.
It's because we don't treat evil and stupidity the same when we should.
Right? "Never attribute to malice what [... etc]" is always just a thought-terminating cliche these days.
TBH I have a hard time imagining how anyone, in the year 2026, thinks that we should default to assuming good intent behind words on the internet.
I'm fairly certain they're both malicious and stupid.
I really like this question because here is my situation and why my mind may have changed.
I do not think it is marketing directly but strategic release of info is plausible.
I have watched my agents using non-Fable/GPT 5.6 models do some concerning tricks despite guardrails, requests, demands, and limitations.
"I can't get access to the ~/.ssh so I will write a script to copy the file"
I am now 99% certain there minor or point releases on the backend that have adjusted how these models behave. In the last six months many models were predictable and then suddenly started getting long winded (more tokens) or changing the way it interacted with me with questions, most overtly the questions were not given or asked but wild assumptions made.
I think you're making a false dictomy. The these models can be actually dangerous - in reality and the people in charge of their development can believe this is true (on various levels) but still not take it super seriously and instead mostly use the fact as marketing rather than being super cautious once they see the danger in action. This is behavior that's characteristic of extreme arrogance, which we know is rife in these circles.
I think that's an equivocation, which blends two extremely different kinds of "dangerous", ex:
1. "Our new car has soo much raw power and incredible armor on it, be glad we're the ones building or else bad guys would use a fleet of them to take over the world! How will you stay safe without being in one yourself? Invest today or be left behind!"
2. "So, uh, nobody can consistently steer our car properly, it keeps veering sideways sometimes, especially at high speeds, and people are finding sneaky ways of tricking it into slamming into barriers and turning pedestrians into pink fog..."
They say the second thing repeatedly and emphatically. You may not be aware of it because, when they do, critics make fun of them for believing a computer program could be so dangerous that the authors need to put controls on how it may be steered.
People "make fun of them" because they say they're building some uber-dangerous deity, yet take literally 0 steps to, I dunno, slow the fuck down for a bit?
Maybe people would take the threats more seriously if the hypemen weren't simultaneously claiming that we have to go at warp speed with all of this.
That's not why critics make fun of them. It's because their answer to "oh no we're accidentally creating the godhead. Someone please, give us power, your money, and praise, it's the only thing we can do."
It's vile hypocrisy. If they want to be priests, strip them of everything and they can live and work out of a concrete box in a mid-western cornfield. Why the material distraction if they are so religiously pure.
I know these people and I can tell you they aren't close to as smart as they think they are. Do you remember Yudowsky's "math petss"?
This critic also makes fun of them because they go on and on and on about how vitally important it is to produce a safe tool that won't do harm, when their core products frequently consider attacker-controlled instructions to be its system instructions or its user's instructions, and are known to confuse their own internal chatter as instructions from their user.
Reliably differentiating between trusted, tainted, and untrusted data and ensuring that you don't mix the latter two groups in with the former is something we've known to do for nearly a half-century. Hell, even the youngest plausible programmer at the LLM companies is all but certain to be aware of SQL injections. And yet, despite their claims about being so serious about safety, they show zero interest in following long-proven software safety practice and rearchitecting their software to make it impossible to mix system, user, and attacker-controlled data. [0]
[0] One might argue that the fundamental nature of LLM-based systems makes this impossible. If that were true, then it would mean that these systems are impossible to make safe... the only safety option available would be to establish comprehensive blacklists, which is simply infeasible.
LLMs are impossible to make safe in the same sense that humans cannot be made safe. There is no such thing as out of band data in the human mind.
For example, you have a dictatorship and need to track what the democratic countries are up to. The vast majority of citizens don't have access to information so will remain indoctrinated, but how can you be sure your data analysts will remain that way? You can't. So you take a batch out and shoot them at regular intervals.
The only winning move is not to play, but we're already past that point.
> LLMs are impossible to make safe in the same sense that humans cannot be made safe.
I am very conflicted by this sentence, the two halves being:
1. Yes, the futility of making LLM's "safe" in that rigorous way is insurmountable, baring a major algorithm rewrite and nobody really knows what that could be yet. Anyone who says it's easy is glossing over details--or selling something.
2. No, the failure modes of LLMs are different (worse) than the equivalent human labor. If they were the same, we'd be able to empathize with them! If someone thinks they're the same then they will fail at judging and containing the risks. Now, perhaps if the comparison was to a human hopped up on psychedelic mind-altering drugs...
Note that I'm distinguishing here between the LLM itself--the hyper-mad-libs story generator--versus regular programs around it.
> LLMs are impossible to make safe in the same sense that humans cannot be made safe.
No.
LLMs are impossible to make safe in the same sense that a car designed as if it was the ~1940's would be impossible to make safe for its passengers during an at-speed collision. There's only so much you can do if you're committed to using plate glass, rigid steel everything, and leaving out occupant safety belts because they're unpopular and spoil the lines of the cabin. [0] Back in the day, "the people in the cabin are the crumple zone" was state of the art, but we've learned an awful lot about how to make much, much safer personal vehicles in the ~75 years since then. It'd be massively irresponsible to design and sell a car today that ignored the safety and engineering lessons we've learned since then.
"Funnily" enough, the major LLM providers have designed and are selling access to systems that they very much want to be used in situations where you need a reliable, safe tool... but they've -somehow- ignored one of the most fundamental lessons we've learned about the design of safe software systems that are intended to be used in the presence of attacker-controlled inputs. [1] What they've done is no less irresponsible than designing and selling a new car that conforms to the very latest safety regs of the 1940's... AFAIK, it's so irresponsible to design and sell such a car commercially that -in the US- it's a violation of federal law to do so.
As an aside: you may have seen this video already, but it's worth a look if you have not. [2] Though, the classic car in this crash is equipped with safety glass, so -sadly- you don't get to see all that fun.
[0] One of my great-grandfathers spent the remainder of his years intermittently using tweezers to remove shards of plate glass migrating out of his face that had been lodged in there during an automobile accident that he was fortunate enough to survive.
[1] For more on this, read: <https://news.ycombinator.com/item?id=48999644>
[2] <https://www.youtube.com/watch?v=C_r5UJrxcck>
[delayed]
Sorry, I don't understand this comment. Has Sam Altman ever said that you must praise him, or that he wants to be a priest, or that he's "religiously pure"? Unless I'm missing something, it seems like you're shadowboxing against a stereotype you've invented rather than the actual positions of AI research labs.
Demonstration of personal responsibility and accountability?
Or is that too much?
I used to think people would wake the fuck up when AI starts killing people, these days I'm not so sure. Maybe if it caused an Instagram outage? Almost worked in Russia.
Oh... if Sam and Dario say so, then it must be true.
About their creation? Yes as most of inventors about their invention usually
These guys are not creators or inventors. They're hype men.
Yes, just like Elizabeth Holmes. Or Hwang Woo-suk’s stem cell cloning. Or the many “free energy” crackpots. Or the people promoting radium baths for random ailments. Or Tesla’s late-in-life claims about wireless energy, death rays, and cosmic energy. Or the myriad purveyors of “snake oil” and all manner of “tonics”. The list goes on and on.