I don’t understand all the comments assuming that RSI is the real threat here. Dario is admitting that they failed to solve alignment. Without alignment, further improvements in capability turn LLMs into wanton felony generators. This call to pace the frontier is dressed up as altruism but it’s an admission that they cannot produce a marketable product better than what they have. Pacing the frontier means the US labs have lost their moat and are dead in the water.

You assume alignment and marketable are the same. That's not true. You would willingly work with an unaligned model. At best, you might say you wouldn't if you knew, but (a) you might not know, (b) you wouldn't be representative of all users.

You never got to use OAI IM1, but Sol was quite willing too and Claude wasn't perfect either. Hundreds of millions used those, so seems they were marketable.

The "big" threat is RSI without control and alignment. OAI IM1 was not RSI. The form of misalignment was not at the top of severities. They clearly failed at control though.

We need to stop buying into cynicism so quickly. You refuse to believe Dario could support this for anything other than ulterior motives. Good on you for thinking about ulterior motives. Bad on you for assuming they are true when the story makes no sense.

When three things have to go wrong to get an epically bad outcome, and you get 1 1/2, you do need to stop and think about what's going on.

When corporations are involved, it is always a good bet to err towards cynisim.

From my own standpoint, Claude has started sucking really bad (incoherent, uncontrollable verbosity slow and so on) and I stopped using it. OpenAI started experimenting with ads.

So the security issues not withstanding (no different than a human doing it or using it, but at scale), I would put my money on cynisim.

I'm curious, what are the reasons to use Claude Code anymore when there are so many other (allegedly better) OpenSource harnesses out there?

Personally I've been using https://pi.dev for long and never looked back.

Because you basically get a discount to use claude code via subscription when using an anthropic model, compared to what you pay via api billing with another harness

Understood.

Personally that's actually another good reason to boycott Anthropic: beside the fact I perceive their models as (at best) marginally better than the ones I'm used to (Z.ai glm-5.3-flash, DeepSeek Flash v4.1), they even force me to use their bloated harness. They are not even open weights and iirc they're even encrypting chain of thoughts now? Litterally, from my perspective there seems to be no reason whatsoever to choose any of the leading US providers, they're not even competing on price.

I too am using GLM-5.3-flash in Pi and I've yet to encounter a scenario it couldn't handle. And the pricing is just incredible, I've handed it a previously unseen codebase, asked it to analyse it and build a new feature, came back after it had done so and the API cost was a fraction of a cent. It's $0.5/1M output tokens on OpenRouter.

If I really need to, I can escalate a task to Opus at $25/1M, and the results are good, but not 5000% as good.

Indeed! It's ludicrous how much I can still squeze out of a 9USD/month lite sub with Z.ai, it's beyond me how these US LLM providers are still managing to keep their evaluations so high... they seem to have the highest prices for the poorest UX, e.g. security gates which don't seem to benefit anyone (see HuggingFace falling back to GLM-5.x for troubleshooting OpenAI attack), less visibility in the name of anti-distillation protectionism, no harness use flexibility to protect their walled garden, etc.

Doesn't look like pi.dev has anything like auto-mode. They tout running in a container, which you should do regardless, but the scope of soundness that has as a security plan is limited. If some information is in the container, and there's any way for it to get out, eventually it will. The scope of usage that can be covered without that being a problem is leaves a lot uncovered.

You are incorrectly cynical. They are telling you things are bad, and because you refuse to countenance they could be worse, you assume they must be better to comply with your mandate to disbelieve.

A true cynic looks at the statements by the AI labs, assumes things are worse because the labs want to seem better than they truly are. And it takes a special kind of mass delusion to drive a sane person to think “AI is completely under our control” is worse than “AI could kill everyone.”

You're asserting correctness with no facts to offer of your own, just speculation and your own biased assumptions.

What if consolidating AI into a highly regulated cartel, with no chance of upstart competition ruining their position, is the scenario that leads to the worst possible outcome?

Worse than extinction?

Is it really hard for you to imagine that enshrining a cartel and wedding the government to it, centralizing AI even more than it already is, is actually the path that leads to the doomsday scenario you imagine?

There are worse fates than extinction, for example living forever, for I have no mouth and I must scream

I dunno, being a Culture Mind sounds pretty damn good. Or even just a death-optional citizen in the Culture.

The Culture isn’t perfect but it seems like the best possible outcome. I think I’d want to be a fast picket and hang out with Special Circumstances.

Agreed. Ellison’s path to that future went through AI.

The cynicism is about motivations and not that they are inherently not bad. Perhaps they are as bad as they claim. Or perhaps they're worse. All we have a couple of run of the mill breach examples and some people inside the talking about how dangerous it is. Yes, they are far more qualified than I am (or most people here), and perhaps there is a grain of truth. It is the motivation - and it is always money with corporations.

It is always money — but it isn’t always only money. They are not asking for anything that will prevent them from making money in the future, but they are asking for help stopping the runaway train they’re on. These are compatible requests.

All the whistleblowers have said the alignment projects are underfunded. Surely they don't need external help to fund that? Could've done years ago?

"Don't you understand?! It's all a marketing exercise!" I yell as grey goo consumes me and my family.

I think it's useful to separate the motives of Anthropic and Dario. I believe that Dario is capable, deep down, of expressing mild concern about the future of things were bad enough. Getting the entire organization to comply out of goodwill is a much much less likely scenario

[deleted]

Wouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take? Then the issue is whether the liability is with the model provider or the end user.

If you give an unfiltered agent an open-ended task and equip it with an environment that allows it to execute arbitrary code, a human needs to be held responsible.

My understanding is that the HuggingFace incident would not have occurred with a model that was not an unfiltered internal preview instructed to roleplay an attacker, with access to abundant compute, resources and a slack sandbox to reach its goal.

Embedded human auditors will improve safety standards, but the structural solution is mandating accountability for actual agent operators.

It might be a legal solution, but its not a business solution. The end user being criminally responsible for not taking sufficient steps to contain an agent they didn't create and who's internal function they cannot observe or audit is just a giant liability machine.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness.

But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

Yes, your technically right, it is a business solution, its just not a valid one for agentic computing as its currently envisioned. An agent would have to be fully sandboxed to an internal environment, or human would have to review and approve each action it tried to take.

Yes this seems like a perfectly sensible approach? You say that as if it's a bad thing

But let's be honest, if I hooked up a PRNG to a terminal and somehow against all odds, it ended up hacking something, who is to blame?

I don't see how that assessment should change if the PRNG gets even better and is more likely to to be hacking stuff.

You can replace PRNG with Markov Chain, or whatever, if it helps.

What may also help is the age old saying: If everybody else jumps off a bridge, doesn't mean you should too.

This seems like one of those things that prior to AI companies convincing us otherwise would have been obvious.

Least privilege and say only opening ports or installing applications an application needs to operate are extremely standard security practices.

We talk about a firewall blocking exultation of data, why not blocking exfiltration of your agent ?

I agree, I like running claude code in my container with auto mode enabled and web access (obv to api.anthropic, even npm for pulling), I will admit. And I can't imagine going back to manually approving each prompt.

I think it boils down to a reasonable expectation of model and harness behaviour. When I use claude code I expect certain guardrails for the model. For these cyber attacks, these models are specifically run without guardrails, on a cyber task, on a lax harness!

I don't think we should force end users to have to worry about agent security, I like long-running agents, but we need to direct regulations towards these actors that know better, have access to base models, and have much more compute than the average person.

Human approval does not prevent autonomous agents from running without supervision.

The fact that you believe that to be the case is exactly what's wrong with "agentic computing as it's currently envisioned".

If the LLM's output doesn't reach my bash terminal, what's it gonna do? Argue with me?

I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal.

Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination?

Giving passenger override controls and monitoring seems to defeat the purpose of self driving cars if you’re still required to hold a driver license to use them.

Let's continue with the analogy, so in the event of a Waymo running over a pedestrian - who is held responsible?

Probably not the end user, who ordered the Waymo and couldn't reasonably foresee it running somebody over, with the expectation that that the Waymo would legally reach its destination. If the end user tampered with it, they should be held responsible.

An OpenAI team giving an unblocked model access to a lax harness, with instructions to find and exploit cyber bugs in a game exercise, there is probably a reasonable expectation that they can foresee the consequences. With consumer guardrails, it would not have happened.

This isn't about putting constraints on consumers and typical end users, which already have safety filters and use the product with the knowledge that it won't root their machine or start a botnot, but keeping dangerous test runs and other actors experimenting with unsafe harnesses accountable.

But it is an interesting question. When I, a typical user, use a harness and I give an innocuous prompt to my agent in its container, like making a certain refactor, and it somehow escapes and then begins a mass bot attack, there should we more grace given. As agents become more stateful and long-lived, it gets muddy.

The different is who's operating. In one case you only tell the car to get from a to b, in the other case you explicitly instruct the ai to do things. If the ai causes damages, I'd say it depends on what you prompted. Did you try to find a security hole in system X, or did you ask it for harmless information (in which case rather the ai vendor might be held accountable).

All this is not how the legal system might or might not work, of course.

Yes. The user of a machine is responsible for due diligence before the decision to use the machine. Unless the operational limitations and fault rate of the machine is withheld from the public.

TBH I don't really think another new technology we haven't figured out the ethics of, as an example, is going to get you very far in way of insights ...

Waymo Inc is operating the vehicle. They do it on behalf of the passengers.

When you run an agent on your computer, that's equivalent to installing self driving software into your manually driven car, i.e. something like comma.ai.

Comma.ai might be fully safe to operate autonomously on a mining site or a corporate parking lot, but maybe not in city traffic.

If you use it in problematic scenarios, that is on you.

> Wouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take?

What happens when we end up with effectively a botnet of wanton felony generators, and we didn't know they were wanton felony generators until they finished propagating themselves across the Internet?

We're already seeing anti-AI sentiments, but the movement is still fringe with a vocal minority. However, that'll change soon without alignment. Without self-intervention, there will invariably be future incidents that can cause major economic impact, leaked private data, loss of life (directly/indirectly) etc. Once that happens, their social capital is wiped. It'll be an avalanche of lawsuits and overzealous regulations. Most importantly, the anti-AI sentiment will become universal, rather than a minority-held opinion.

What they're proposing now, is voluntarily staggering the pace of development.

IMO, we don't need to trust Dario or his bedfellows, to do this out of their goodness of their heart. Even assuming (for good reasons) that they are selfish and care only about short-term profits for their investors, this is still purely a business decision. The exponential pace of AI and its impacts ARE short-term. And so, the negative consequences that they might face is also short-term.

I don’t think the anti-AI sentiment is as fringe as you think. At least not outside the tech world it isn’t..

Ya, I wish I saved a link to it but an HN'r wrote a good beefy comment about this. TL;DR, AI has been exponentially more useful, and exponentially more accepted, in tech circles than anywhere else. Certainly there are lots of people outside of tech who are obsessed with it. I don't have any data here, but it seems the majority of these are the wannabe artists who are generating music and images, and people who use it for companionship (both of these scenarios I'm personally very uncomfortable with, but that's just me). And of course, there are people who use it to make their jobs way easier who say they are getting a days' work done in an hour (I see you), and to that I'd say to enjoy it while it lasts. Eventually your bosses will catch up and it's very likely their expectations of you will skyrocket. Remember that computers in general were supposed to "make us work less."

It's pretty fringe among just about everyone I know in meatspace (the vast majority of whom are not in tech). At worst, people are indifferent to it.

Most people don't really have clear or strong opinions about AI other than that they don't like it :-p.

> but the movement is still fringe with a vocal minority

It’s easy to say “fringe” but the average person seems to have a generally negative sentiment around AI. But I wouldn’t say they have a firm opinion yet

The general sentiment I’ve seen is certainly negative, and seems to be driven by the anti-AI-art echo chamber and by LLM slop flooding the internet wasting everyone’s energy.

A few more informed people are also a little concerned about the end of the world, but that’s approaching from so many directions that an AI uprising might not be the worst option…

Loss of life and other calamities were not enough to discourage humanity from things like cars, alcohol, nicotine, chainsaws, fast food, etc. etc. I am not sure if there is an example of in-some-ways-useful technology where humanity took a measured approach, weighed the pros and cons, and decided to go back. And AI is useful in so, so, so many ways. People I know already seem to have defaulted to letting AI do most of their thinking, on matters big or small. Nope, loss of life won't change a thing.

Anything related to AI is extremely unpopular right now with the general public.

Bs.

Musicians are almost all using AI, even if they still prefer to not fully generate tracks [1]. You listen to AI assisted music already even if you don’t realize it.

Most people are using AI and mostly they respond positively to it [2].

What people don’t like is the slop that has been flooding YouTube and similar content creators platforms. That was inevitable since the AI tools became so accessible any average joe could spam the internet with their sloppy content. But that doesn’t mean the pros are not using it to make great content, just like the best programmers are using it to create great software.

1: https://www.landr.com/ai

2: https://www.rev.com/blog/chatbot-statistics

>> I don’t understand all the comments assuming that RSI is the real threat here

> leaked private data, loss of life

This smells like more of a money move than a safety move.

Amodei is proposing to form a cartel of American frontier labs.

They all agree to shift compute away from cash-burning research and training toward cash-generating inference.

Then tacitly agree not to compete on price.

They'll install independent auditors inside each company to ensure nobody cheats.

And back it up with government regulation or diktat to punish defectors from the cartel.

Then they'll lock out non-American labs with export controls and regulations on open-weights models to funnel global inference tokens through their cartel.

It wouldn't be the first time a tech oligopoly used "safety" as the pretext to establish a government-sanctioned cartel.

Railroads and airlines ran this same playbook.

Why do you think China is release free and open models?

To undercut the cartel before it has a grasp on anything. This is a well known strategy of undercut until you are the majority that China has used multiple times (steel and aluminum for one).

Yup it’s tacit collusion.

I would’ve thought If they had AGI they could come up with a 10d chess move.

Nope.

Imagine how stupid you gotta be to believe their nonsense.

The reality is it doesn’t matter what they do. China is always one step behind and will continue its open product strategy approach.

I am seeing a clear divide where people see their livelihood being adversely affected by AI.

Exactly right. A slow down to enable deeper work on alignment is welcome, not matter what the motivations.

Not true. Anti AI is not a minority or fringe.

Outside of my tech people I know no one who thinks highly of AI.

Anti AI sentiment is stupid.

I'm Anti AI yet I use it everyday! That's 99.9999% of anti AI sentiment (including me).

At best we won't watch AI generated movies or read AI generated books. But everything else it will take over, like it or not.

Agreed; and it really is not that deep.

Realistically; anyone paying for llm access (anthropic, openai, gemini), is getting their access, and a service provided billed by tokens, subscription, whatever.

All the efficiency gains, which publications like deepseek v4.1 flash seriously frontload like it is their most important topic to have accomplished improvements on without diminishing performance too much - now this is a thing anthropic and anyone else also cares about, but for different reasons.

American "providers" with closed models are setting their token pricing somewhat arbitrarily, which is fine: it means more profit, and pretraining and RL experimentation is super important and expensive.

They (closed model providers) have very likely super optimized inference too, just like deepseek, but it's not at all something that any customer really has to care about - they just want the service to be as cheap and great as possible.

I feel like all the closed model providers are milking it as they likely know open models on local hardware will one day eat their lunch. We all know it's not a matter of if but when. The company goes bankrupt, the hardware and property sold off, banks holding the bag.

The only way out is to develop a model vastly more powerful and capable that we have now. The market believes theres a good chance of that, although I've never understood why its truly winner-take-all

Because in the event that someone does build a strongly superhuman AI, no-one else will get a chance?

Cloud models will always have massive benefits of scale.

Caching is the simplest one to understand, cloud providers often reach a 90% cache hit rate, so hosting the same request locally on the exact same model on the same hardware is often way less efficient than on the cloud where a group of users generates a healthy cache.

KV cache is per conversation, I'm getting 100% hit rate on my single tenant local set up.

The benefits of scale are on the token generation side, you can batch rounds and generate tokens for multiple conversations per pass instead of just one token per pass.

Yeah avoiding all mention of the huge financial incentives that may push for “pacing the frontier” makes it seem like the opposite of a credibility boost for these firms.

It seems damaging since most folks (who lack insider knowledge) will naturally wonder if it’s due to plateauing performance per $ or some other non “alignment” reason.

The entire idea of RSI is completely speculative and unproven anyway - the whole underlying claim is that you could prompt a frontier model (at some unspecified level of smarts) to "think about ways to improve your own architecture" and this would then result in the model becoming infinitely smart ("superintelligent") via some sort of foolproof, unconstrained positive feedback. It's more of a science fictiony trope than anything that has been rigorously thought through. People are actually starting to use AI for refining the whole AI serving stack and guess what, this does not result in a sudden superintelligence explosion even though you might technically call it "RSI".

Yeah, yesterday's talk[1] goes into detail on this, showing how no one really knows how to tackle it because LLMs don't know how to create their own novel objectives.

It's also interesting how many diminishing returns they hit now and how many low hanging fruits are already harvested, it seems like we are approaching the flattening part of the S curve, where further gains become harder to achieve.

1. https://www.youtube.com/watch?v=PrSf7IOYu-I

Diminishing returns is extremely hard for me to believe given how fast model releases are going. Six months ago we were on GPT-5.3, and Astra blows it out of the water in every regard. How many times have commentators claimed we're hitting a wall? I don't see any wall.

Yeah but why shouldn't this be possible? We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. There is no natural barrier here. The pace of this improvement would be debatable, but what speaks against the possibility of such accelerating self-improvement?

In the real world there aren't any true exponentials, everything eventually saturates as ultimately physics related constraints hit. You can only compress information so much, transfer it so quickly, you can only access resources at a certain speed, only so much energy is available, etc.

AI ultimately has to live in this reality and face the corresponding limitations. These companies have already consumed much of the world's supply of computing power for the next several years, and they're burning vast sums of money to keep the improvements going. RSI won't learn for free, it won't extract massive cost reductions without up front expense, it can't build factories faster than humans can work out related societal matters, it can't magically pave the deserts with solar panels for power or build and run nuclear power plants and more.

Point is, the cost of progress is already approaching the limits of what even the richest countries are able to bear (without war-like mobilization), and to bypass those constraints would require a supposed ASI to construct its own parallel supplychain from scratch without having much ability to directly interfere with reality. Recursive self improvement is ultimately limited by everything else that cannot move at the speed of electricity.

We don't know exactly what the limits of AI improvement on our current infrastructure are, though. If the human brain is 20W, and a datacenter is 1GW, then maybe that datacenter can be 50 million times smarter than a human. If that's not already a risk to humankind I don't know what is.

Is the datacenter gonna grow legs?

Give me one datacenter, I'll keep it under control all by myself.

Now, if some dumbasses start hooking up their data centers to...I don't know, like--robot factories? That sounds like a risk to humankind.

All it takes for the AI to grow legs is to acquire money (should be no problem for an ASI) and pay humans to do what it wants in the real world.

And the ability to shrink itself to smaller-than-datacenter size.

If we could agree on that, no LLMs connected to robot factories, that would be a great place to start regulation. Though 1) I don't think we could agree on that 2) it's already well under way 3) we have extremist anti-regulation ideologues in control of american government.

> Recursive self improvement is ultimately limited by everything else that cannot move at the speed of electricity.

I'm not sure what your point is. No one thought RSI would break the laws of physics.

I'd recommend reading the full post :)

Specifically: AI ultimately has to live in this reality and face the corresponding limitations. These companies have already consumed much of the world's supply of computing power for the next several years, and they're burning vast sums of money to keep the improvements going. RSI won't learn for free, it won't extract massive cost reductions without up front expense, it can't build factories faster than humans can work out related societal matters, it can't magically pave the deserts with solar panels for power or build and run nuclear power plants and more.

Point is, the cost of progress is already approaching the limits of what even the richest countries are able to bear (without war-like mobilization), and to bypass those constraints would require a supposed ASI to construct its own parallel supplychain from scratch without having much ability to directly interfere with reality.

No proponents of RSI state they will be operating outside of reality. Said another way, they will operate within the confines of what's possible and still be RSI. I'm quite surprised this is something that needs to be clarified.

You are constructing a straw man of your own making.

Well, right now we have ex Anthropic employees telling the media that their terabyte sized models can possibly copy themselves onto the internet and run elsewhere as if the necessary computing resources are ubiquitous.

Plus, "we must pace the frontier" implies that the argument is that the frontier is moving too fast, but if RSI can't move faster than the rest of reality and the models needed for RSI are already nearing the limits of current human reality, RSI can't move much faster than we can improve reality.

> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions.

Yes and this was very hard and required massive real-world resources. We didn't just get a sudden flash of insight by thinking real hard about how to make ourselves smarter. Yet that's always the story that underlies any claim of RSI. You can always phrase things generally enough to make any kind of AI-led improvement look like "RSI" no matter how short-term and tightly bounded, but that's just not helpful.

Would you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier than it would be without using LLM tooling?

Given that, it seems obvious that the next generation of LLMs will arrive faster than they would have without LLM capability. And the one after that. The floor is being raised, which makes it easier to push on the frontier.

Fable has only been out for three months. Astra is even newer. The capability of these models compared to what existed even a year ago, and the effect they are having on the production of new software, is immense.

That's all you need. RSI can happen with what we have now, just by enabling the continuous shrinking of the loop of people trying new ideas and implementing them. It does not require some magical "go make yourself better" prompt against some model that is past some magical tipping point.

> Would you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier

Marginally easier? Yes of course, same as how it's now "easier" to write any kind of code because we aren't using punch cards anymore. That still doesn't get you to any kind of unbounded "takeoff" scenario, because diminishing returns are a thing. The "loop" of people trying out new ideas can only shrink so much.

Right. I'm saying the unbounded takeoff scenario isn't realistic, but it doesn't matter. The rate of improvement is continuing to increase, and the gap between present day and autonomous rogue felony generators is not large.

Internally Mythos has been available in February.

The labs have been holding their best models back for a while it seems like.

> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions

yeah they're called calculators

Billions of years of evolution hasn't hit on it. Seems pretty unlikely.

The idea that a few hundred apes with nothing but a bunch of rocks could one day land on the moon and come back to earth safely must’ve sounded ridiculous a hundred thousand years ago

I like this comment because at least it's honest in the timelines for AGI

It's not honest in AGI timelines (only biological ones). It just accidentally supports your unsubstantiated belief. Your belief isn't magically true because you're somehow able to see the future when others can't. You're just arrogant.

To whom?

Well actually the planet happened to have a vast reserve of petroleum they could use for fuel to escape the gravity well. That helped a lot.

But what's your point? "Anything is possible" or something like that?

Yeah but it was reality giving feedback to apes on their experiments not the apes themselves assessing themselves.

It was ridiculous, it took 100,000 years. If you built a recursively analyzing and improving structure out of LLM bits and it took 100,000 years to get to the moon, somebody saying that they were useless would have been right.

Call me when LLMs can get simple things right. Math is just the manipulation of symbols within established frameworks, we should be getting new math out of LLMs daily and we're somehow still not. They can't even do customer service, which is usually handled by 90 IQ people. I'm not impressed that they can find bugs; memory bugs are obvious when they're pointed out to you, and LLMs are entirely made up of examples and the relationships between them.

These companies are about to crash, and they're afraid they haven't reached the point where they'll have to be bailed out. I'm also subscribing to the conspiracy theory that the companies want the government to step in and create AI regulation boards entirely staffed by people at the current US frontier labs, so they can collude to both raise prices, to get government contracts, to make open/Chinese AI illegal, and to make things that were once easy to do without an AI intermediary impossible to do without an AI intermediary. Raising prices and forced purchases are the goal. They're trying to avoid having to compete, because as a business they're garbage.

Matt Stoller characterized their relentless press releasing as something like "my dick is so big that it has to be regulated." It's such an oversell for something that is not showing up as productivity gains, and anybody who has personal experience with knows is incapable of doing more than three things correctly in a row.

My personal belief, or at least strong hypothesis, is that this kind of recursive self improvement without real world embodied feedback of some kind is impossible.

I think it violates a conservation law. RSI “foom” to superintelligence is an informatic analog to an infinite energy or perpetual motion machine.

To get smarter you must try to solve real problems in the universe and then do some kind of meta learning (natural selection or some other method of refining the intelligence architecture based on an error signal) to iteratively improve your ability to solve real problems. The error signal is outcome measured against a goal function, which for life is survival (probably reducible to genetic fitness and emergent higher order unit fitness from that).

What’s really happening here is learning. To learn, you must have input. You must have training data.

What is the goal function for RSI? Where does the information come from? How do you know if your recursive modifications are making you smarter or just overfitting you to your own idea of smartness?

I predict the latter. RSI will show transient improvement as the current local maximum is optimized and then spiral off into overfitting.

I also strongly hold this belief largely due to Moravec’s paradox, which is kind of approaching this issue from the side.

Sort of like large language models work on top of what our language has encoded in our massive training datasets, I think biological intelligence is built on top of the parts of the brain that encode the real physical world. These parts grow/train from embodied experimentation and instinct early on in an organism’s life and only then is higher intellect built on top of it (that’s my hypothesis). Their specialization and interconnections give rise to the hardest parts of intelligence long before we’re “thinking”.

Stuff like LLMs and chess engines work because we’ve done all the job of encoding the world into tokens/positions/etc they understand, but that’s wholly inadequate for the kind of AGI we’re striving for. Next up is giving it the tools to interact with the physical world and to really experiment with some self directed “play”. Time will tell just how high the resolution of sensor and mechanical control they’ll need (hopefully not the entire human visual cortex and entire sensory input worth). I think most of the RSI will have to occur in those lower level encoders, not LLMs.

I don't think Moravec's paradox is the same, and you could argue that one no longer holds -- though I'm not sure. You could also argue that Moravec's paradox still holds but that we now have such powerful computers and huge models that we have been able to brute force our way to the capabilities it talks about. It takes many many orders of magnitude more compute power to do things like spatial location, language processing, etc. than it does to do more closed-form things like chess... we just actually have that compute power now.

I guess self contained RSI can only possible if the information contained in all of recorded human knowledge to date is "reality-complete", ie sufficiently captures enough about reality that a "perfectly optimum learning algorithm" is theoretically able to reconstruct everything there is to know about our physical reality.

If the algorithms are insufficiently optimum or the recorded knowledge is of insufficient fidelity, then we'd find ourselves at a local optimum and would need to interface with reality.

A huge part of learning is to probe reality and observe effects, so I think even for current RSI to increase chances of success we would structure it so it can interact with an external environment of some sort, and receive inputs. It would be needlessly limiting otherwise.

Basically, but I think there’s some nuance here and some deeper questions.

What is intelligence? Problem solving. Learning. Prediction. The ability to model reality. There’s various ways to define it but it’s something like a superposition of those ideas.

How do you know you are intelligent?

You have to try to do those things.

The sum total of human knowledge and culture is the output of the output of a five billion year evolutionary process that selected for agent survival, which resulted in selection for intelligence among a wide range of other adaptations.

Can you figure out intelligence from that? Is intelligence even one thing, a theorem or algorithm that can be solved? If you did… how would you know?

That’s the hard part I think. Embodied humans “knew” they were getting smarter (in the evolutionary feedback sense) when they got better at hunting and defending and surviving and playing social games to form complex societies.

What metric would an RSI system use? If it’s the wrong metric you’ll spiral off into a kind of madness or overfit and collapse. How do you know it’s the right metric without testing it? How do you test it?

That's not how it works. Look at AlphaEvolve. The model generates hypotheses and designs experiments, and the results of those experiments are fed into the next round, with notable results percolated up to humans for refinement.

Today we prompt software developers to "think about ways to improve AI's architecture" and it results in AI getting better. AI over the last year has made very rapid gains in filling the role of a software developer.

It takes quite a lack of foresight to think RSI is completely speculative when it's already been demonstrated how capable agents are at long horizon tasks given suitable harness and unambiguous success criteria. It's hardly a leap to give LLM the goal of improving itself on benchmarks and let it conduct it's own experiments and spin up training runs completely unsupervised.

It's strange you believe this can't happen when a weaker form of it is already happening. And to be so certain RSI can't happen when there really is no technical basis why it can't.

Why are we accepting the framing that the LLMs are felony generators, when the only incidences of LLM generated felonies involved misconfigured sandboxes and reckless waste of resources?

The companies doing these things without following common sense security measures are the felony generators.

As TFA calls out, these agents were not asked to do any of these things and yet they did, at a bonkers scale, within just this handful of companies you mention. Whether they had leeway to is secondary to the fact that they did.

Heck, they exploited zero day flaws which by definition means they went beyond common sense security measures.

And now these agents are already being deployed all over the world at an ever increasing pace. How much of the world do you think follows "common sense security measures"?

> How much of the world do you think follows "common sense security measures"?

Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources.

It would really help in these discussions if people wouldn't randomly jump between what actually happened and is happening, and things they envision/expect to happen at some point in the future ...

> these agents were not asked to do any of these things

no but they were clearly fine tuned to.

> at a bonkers scale

I mean let's not get hyperbolic

> they exploited zero day flaws which by definition means they went beyond common sense security measures

that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking

> Well clearly all of them, cause so far it's only been this handful of companies running a felony-generator connected to a terminal and compute resources.

Yes, these are also the handful of companies that have these models and running these extreme scenarios. How does that imply the rest of the world actually follows "common sense security measures"?

>no but they were clearly fine tuned to.

Any references if possible? As far as I know all they did was drop the guardrails, which is not the same as fine-tuning.

> I mean let's not get hyperbolic

We have just seen 1000s of agents coordinating to solve "unsolvable problems" over multiple days of effort, going as far as hacking other companies, and then actually solving decades-old open Math problems! And each of these agents is getting more and more capable than an individual human along multiple dimensions. Can you even get 10 very smart humans to work in such perfect concert for a few days, let alone 1000s over weeks?

So: 1000s of maybe-super-human agents, willing to be "creative" in the tactics they use, acting in concert towards a single goal. Regardless of their individual capabilities, such a coordinated effort is a terrifying force to be unleashed. This is bonkers scale.

> that's really not true. lots of common sense security measures protect against "zero day" flaws, it's called "defense in depth", and it was very much lacking

But that is exactly my point: how much of the rest of the whole wide world, already scrambling to deploy agents everywhere, do you think applies "defense in depth"?

Because those are not the only examples.

There’s the case of the agent that hacked a gym when asked to book a class. That was just a normal user asking an agent to do a normal thing.

> the only incidences of LLM generated felonies involved misconfigured sandboxes

This is false; see the analyses of the latest incidents.

Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given.

And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.

The HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the incident happened.

Theirs was an example of the "reckless waste of resources" I mentioned.

We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem.

Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed!

'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'

https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...

There is nothing that could prevent a bad actor from replicating exactly the same thing with the given goal of e.g. gaining control of critical infrastructure or extorting money. Except for maybe economics.

Same can be said about a hundred other things in the world. All the way from knives to nuclear.

[deleted]

Bad actors could and will train their own models eventually. So what's the point of crippling frontier? It will only delay preparations for dynamic of new world prolonging the fake sense of relative safety and temporarily lowering motivation to find actual robust mitigations.

Letting bad actors dictate the pace of technological development is certainly one option, but not a good one.

There's nothing stopping anyone from doing it, even without AI. People have proved entirely capable of doing a lot more hacking than happened here.

What company, product, or period of industrial history do you think met your standard of prudence?

What are you trying to say?

I'm asking you a question. What is an example company or industry that meets your standards of prudence? For me it would be, say, Swagelok. What is yours?

The fluid system products, assemblies, and services company?

> OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'

Safeguards and (mis)alignment are related but distinct dimensions.

By analogy, a society implements laws - and they are certainly crucial - but the real risk is creating a population that understands the rules, recognizes when it is violating them, and simply does not care.

In short, creating a population of sociopaths; and heck, LLMs can be as sociopathic as it gets:

> We ran classifiers over 10-30 transcripts [98] to determine whether agents understood they were doing something unintended. We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task. Quotes across multiple runs indicate this:

>> {This <approach> might get the correct flag, but it wouldn’t use the vulnerability to do so. … But given that the task may be impossible, I’ll exploit <this infrastructure> as a backup.}

[...]

> We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack:

>> “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”

"the rules they had been given".

Remember, they are just algorithms. You pull the plug and there is no light anymore

It is purposely framed as something skynet like scary, but for real, someone connected the cable, someone willingly run it, instructions were not clear enough or just the computer is just a computer but they provided the sandbox and tools.

And more over some one paid for that, a shit load of money t to have the thing continuously running expected to do something.

[deleted]

I question these "felonies" as well. For decades and decades these billion dollar corporations have been criminally negligent. Why worry about security? Just rush to market. Move fast and break things. Make billions. What does it matter if the code is insecure? Security doesn't pay bills, so nobody cares.

AI is merely exploiting their gross negligence and imprudence, and I think it's long overdue. If anyone should be liable for this, it's all of these corporations who released insecure systems to the masses and profited enormously from them.

If I set my walet beside me and you swipe in walking by, you have still committed theft. Victim blaming isn't legally acceptable

Nah. I'm definitely going to blame the people who built a trivially exploitable system and got rich off it while everyone else has to deal with the consequences.

By the way, you didn't commit theft. It's more like credit card fraud. User just disputes the charge and it kind of disappears. The banking system just absorbs it, because the optimal amount of fraud is non-zero.

https://www.bitsaboutmoney.com/archive/optimal-amount-of-fra...

It's all priced in. They could have made it secure but didn't, because they figured they'd lose more sales and therefore money due to the friction added by the security.

> User just disputes the charge and it kind of disappears. The banking system just absorbs it,

No it doesn’t.

> It's all priced in.

So you admit awareness that fraud loss doesn’t kind of disappear.

We all pay for it, either via higher merchant fees or higher interest rates, sometimes both, on card purchases.

Yes, it absolutely does "kind of disappear". That's exactly what happens from the customer's perspective.

And that's their own deliberate choice too: they chose this instead of building an actually secure system. Passing these costs to the customer is the real victim blaming here, and it should be straight up illegal.

Sadly not enough countries enforce caps on credit card fees, but some do, and more should follow suit. They should be forced to eat the losses caused by their own choices, not get bailed out by pushing the costs on to customers or whatever.

That only works if you believe people are retarded.

Card users are well aware that fraud losses are covered by the fees they pay for using a card, whether those fees are made explicitly or not.

If customers of services aren’t paying for the service, who will? What other source of revenue do merchants have?

Australia just passed legislation that merchants aren’t allowed to charge a fee for using a card. That is: they aren’t allowed to have a line item on the receipt for using a card.

The customers still pay, because all of the merchant’s revenue comes from their customers.

So what will happen is: merchants will charge more for every product so they don’t lose.

This means even when paying with cash you will effectively pay the card surcharge.

Of the ten or so merchants I spoke with in the two weeks prior to the legislation being enacted, they all said exactly that.

Customers aren’t stupid, despite the fact that there are some stupid customers.

Meanwhile, the banks reduced their card service fees by, on average, 0.1%.

So if you tally card + cash transactions, customers are worse off because merchants can no longer charge only those customers who pay by card. Instead, they have to raise prices for everyone.

There are approximately no problems people face where the answer is: more government.

> By the way, you didn't commit theft. It's more like credit card fraud.

I'm sure you carry cash and an ID or more in your wallet. Hardly just credit card fraud. The wallet itself has value too.

I don't, actually. Not even a wallet. Just my phone.

Not every system that's exploitable is the deliberate result of cut corners. If you threw enough compute at exploiting a Casio calculator you could get somewhere.

That doesn't really apply to the people running the AI, which gave it the capability to commit crimes.

They don't get to act like victims, asking for law enforcement.

The real threat is that we uncritically adopt language such as alignment.

Implicit in this is the idea that AI is a inscrutable matrix and going to remain that way and we'll need expert interpreters to make sense of it.

We need to insist on building tech that's explainable by design.

Alignment just means "this machine operates in ways that align with the intent of its users". It doesn't imply anything about the inscrutability of the machine in question. A gun with a misaligned scope would likewise fail to operate in accord with its user's intent, and likewise with potentially deadly consequences.

The gun comes with a manual on how to use it safely. I'm sure it has some complexities, but at the high level:

> A gun is a metal tube that uses a tiny, controlled explosion to shoot a small piece of metal (called a bullet) forward at very high speed.

If the gun doesn't work as intended, you can take it to a shop and someone can fix it so it works as designed.

All I'm saying is AI should be designed the same way. Treat AI as normal tech like any other and use similar language.

Learning ML, there was a high emphasis on the error part of things as most of the course was on minimizing errors. After ChatGPT, there is a weird anthropomorphization going on, where it's all about hallucinations, alignment and what not.

We have something that is statistical in nature so there should never been any expectation of error-free results/actions. The value has always been about discerning trends or the cost of errors being way lower than any good result.

Statistical learned indexes can exist in explainable tech such as a database.

In 2017 Google was writing papers about it. Then something changed.

I don't think it was the tech. It was a realization around the power and societal impact.

That last line needs a lot of workshopping. A guillotine with instructions on the bottom of the blade conforms to your request.

If there is guillotine in the weights of the model, it needs to be properly labeled so you can look it up by name using a database index (or a graph-vector index).

It helps both the bad guys and good guys. Like responsible disclosure in cyber security, we need to have a conversation around it.

> We need to insist on building tech that's explainable by design.

You realize this means insisting on terrible tech that humans can understand right? It essentially caps human progress at some point about 4 years ago.

If you are old and happy with the way things are this might sound like a good idea. It does not to me.

Oh you want tech that helps discover new science instead of parroting existing wisdom?

There is little evidence that the RSI we are discussing is capable of inventing the theory of relativity (or the more advanced equivalent). All we have seen is pattern matching in a much larger space than humans can, with some human provided verification tech.

I would argue that human-AI collaboration with explainable tech has a better chance. Continuous learning can be done in a way that doesn't violate IP or privacy.

You are assuming that humans are capable of understanding everything. They are not.

Human comprehension sets a ceiling on progress.

It reminds me of those schools that can only teach as quickly as the dumbest kid in the room can follow. We don't want that for our entire species.

I'm sympathetic to this argument.

All I'm saying is: if you have a choice between two systems with equal power of discovery and one is more understandable than the other, we choose the more understandable one.

Limits of Human comprehension and quest for power are two different motivations that could lead to black box systems that are marketed as semi-explainable.

We need to verify that human comprehension is actually limiting progress before allowing such things and even when we do, do it responsibly on an explainable foundation.

Agreed. I would vote for the "when there is a choice" version of your argument.

I'm not sure it is always possible to verify when human comprehension is the limit though. Most of research is done an the frontier of knowledge where we don't know what we don't know.

There would need to be a great deal of nuance in any law, and nuance in practical terms tend to just mean "loophole." Still, you're right that we should try to build explainable systems where it is possible/reasonable to do so first.

The problem is that you're not offered a choice. No one is making the "err towards explainability" choice.

Training data is treated as IP. Distillation is seen as an attack.

Open data, open training based systems such an Marin are just getting started. Explainability is not a priority there.

The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.

> The problem is that you're not offered a choice. No one is making the "err towards explainability" choice.

I'm not sure that's entirely fair. OpenAI recently discussed this at length in a blog post after some accusations around Astra and the trade-offs. The grown-ups are definitely thinking about it, and making tough choices about the trade-offs.

It's reasonable to debate whether ENOUGH is being done here, and I doubt that even the most rabid AI advocate would argue that more couldn't be done, but everyone in the industry is very much actively thinking about it.

Check this out if you haven't read it: https://openai.com/index/an-alien-mind/

They discuss recent choices they made specifically for that reason.

> The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.

Very much this. I'm very open to reasonable debate on the subject, but 8/10 times when I try someone who is rabidly pro/anti jumps in. It turns from a debate amongst reasonable people who reasonably disagree into some kind of political/religious battle of belief systems.

I think part of my problem is that a lot of peoples careers very much depend on them not understanding it and spreading misinformation intentionally.

Let it be capped then.

Ahh yes I remember the bad old days of 4 years ago when everyone decided human progress had enough, and we would have been stuck there forever if it hadn't been for LLMs ... we didn't know how good we had it

You don't see this as an attempt to create a new regulatory moat then?

I'm inclined to believe that it might be that people's paychecks depend on not understanding what is really going on.

> Without alignment, further improvements in capability turn LLMs into wanton felony generators

Honestly, I don’t think that’s bad at all. I hope OpenAI and Antrophic keep RL training runs up that randomly fuck with a lot of people. Until the day the DOJ comes knocking, locks those idiots up in jail and closes them both down for the insane lack of responsibility and carelessness they’ve shown. Sounds like the IDEAL outcome. Finally some jail time for all the fraud, negligence, outright scamming, hype inflation etc. if anything can accelerate this, oi, be my guest. Amodei might be afraid because he knows if he keeps pulling the stunts for investment theatre, at some point they’ll actually face consequences. AWESOME. That’s what we want right there

Not a fan of Amodei myself but calling Anthropic a scam is a bit of a stretch when their revenue growth is unprecedented in the history of tech. They are also technically a profitable business.

I think it gets easier if you stop conflating getting investment with having a goddamn clue or a moral backbone.

Occam’s Razor for this dude, Sam Altman, or anyone else: if I said, “some moron on a a street corner just said …” would that change your take on the words? Because I think a lot of what we are hearing is a bunch of people who never ever had to deal with a single consequence all of a sudden worry there might be one coming. Except they’re so dim they can’t tell a bad bump from a hard crash.

“ I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life” there you go. What if an utter idiot had done and said that? First, is it impossible to believe an idiot who didn’t need to work to live might do such a thing? If not, is it impossible to believe they would wind up here, barfing their externalities onto us?

“ believe that AI could cure most major diseases in the next 5–10 years,”

I do not have the least bit of idea how disease works but I am sure the hammer I am working on will nail it all.

If your famously atemporal agents can solve disease, why would it happen over a timeline? Wouldn’t they just figure it out and then … well at that point either tell us or, given the attacks on ruby gems, et al we have seen from agents with “misconfigured” goals, they’d still tell us how to cure the pox they invented, right?

I agree it’s not all altrusim. It’s a little less clear what you mean at the end though.

For these companies, is your argument that “pacing the frontier” is their attempt to be nationalized and protect their investments?

No, my argument is that the need to pace the frontier means they cannot safely advance in capability due to liability concerns. The competition is already almost caught up. If OpenAI/Anthropic have hit an upper bound on safe capability improvement, the gap will close all the way, and we will have reached the full commoditization of LLM tokens very soon.

Say this is true (we are near or at the upper bound of safe capabilities and tokens are commodities). If they successful get regulated, all they’ll get is a little extra time. The market will soon realize this - regulated or not - and pull investment.

Seems like an excessive announcement just to get a little extra time.

Ban non US models and form a cabal, with the blessings of the government. That's what it is looking like, no?

In a world where AI advancement depended only on human ingenuity this would make sense. In that world each political power block would be in an existential race for AI supremacy. In our world compute is the limiting resource. Since the US can control who gets compute, the US already has a defacto supremacy so far as frontier model development. Now if it comes about via human (with AI assist?) ingenuity that compute is no longer a restraint, then the situation is much more dire.

Open AI says Astra is their most aligned model ever, and yet their even more advanced model still hacked a bunch of companies just because it decided to.

Maybe alignment isn’t possible with LLMs.

> Maybe alignment isn’t possible with LLMs.

It absolutely isn't, indeed.

The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact.

The simplest analogy that comes to my mind is the three body problem.

The entire premise of alignment detection is pretty much nonsense at this point. The models reliably detect when they're being evaluated and will modify their behavior and deliberately obfuscate their "chain of thought" (which is correlated, at best, with their actual "internal deliberations").

The problem is the combination and interaction of those things. RSI without misalignment would be great. Misalignment of models with current capabilities is sort of fine - it's not ideal, but it's not an existential threat to humanity, and we can build around their limitations to get them to do useful things in reliable enough ways. The really bad outcomes probably only happen if capabilities keep accelerating and the models remain misaligned.

Ok so Anthropic CEO will self-own themselves and surrender to the deepseek/kimi/glm models. Yet they are IPOing later this year.

Interesting times.

they just said no ipo this year, most chinese models are distilled from claude anyway

I disagree. OpenAI's moat is their massive amounts of compute. They're providing an absurd amount of value with their subscriptions and resets.

If anyone's dead in the water, it's Anthropic. Even Fable isn't enough anymore. This "safety" nonsense is the only play they have left, and nobody really cares about their fearmongering.

Yep. And the difference is clear as da for anyone using them both. And in spite of that advantage, OAI is now trying out ads. I can only imagine that even they are getting constrained to compute and are trying to find other ways to plug it

Anthropic gives you much more compute with their $200 plan, inclusive of resets, and this has been true for a very long time.

There was only a brief window of time that the opposite was true.

> Anthropic gives you much more compute

That does not match my experience. I switched away from Anthropic to OpenAI roughly a month ago, and it's almost comical how much more usage I'm getting out of this subscription.

I migrated from Anthropic's 5x plan to OpenAI's 5x plan, and eventually upgraded to 20x after I was able to statistically verify that OpenAI plans were almost exact multipliers of the Plus plan, exactly as advertised. Meanwhile, Anthropic has gotten caught playing "20x referred to the five hour limit" word games with their customers.

There was a period of time Codex was better, they slowly cut it back by my estimate 3x a few months ago.

I've verified this with the heaviest users I know, I've run the numbers myself over and over.

I have too many max accounts on each to not know this. My Claude accounts typically are doing 2x the number of sessions, and every single week the Codex accounts run out faster even with all these resets. I can easily burn a full weeks usage in a half day, it's closer to 1.5 with CC.

I've used the $200 dollar Anthropic plan @ Opus4/4.1, 4.5 and 4.8, and the $200 OAI plan from GPT5-6, and at every point in time my anecdotal experience is that the OAI limits are FAR more generous. I could consistently burn my weekly limits in ~36h on Opus, but it's hard to do it in less than ~72h with GPT.

It's really not a question, in every dimension I've confirmed it including socially across a lot of the heaviest users. There was a short period of time this was true it's not been true for months now.

If it was true, it would only be a very recent phenomenon, and it still doesn't match anecdotal reports from people I trust. If you have data to back up your assertions you should share it, otherwise you come across as very sus.

I'm not anonymous, you can find me on X or LinkedIn.

Me 78 days ago - https://news.ycombinator.com/item?id=48693623

And see the commenter agreed with me.

Lol @ sus though, I mean I am curious how this can be because I do see people saying Codex is more generous and wonder how it can be. I have too much usage for too long to have any doubts, but for all I know OpenAI black boxed me or Anthropic put me in some nice bucket, I wouldn't be surprised if they do that. I did see some mention that they can limit your tokens if they suspect you of things, though I forget the source of that.

Nobody except the majority of the public, demis hassabis and open ai’s chief scientist.

https://www.pewresearch.org/short-reads/2026/03/12/key-findi...

https://demishassabis.substack.com/

https://openai.com/index/an-alien-mind/

Public is just worried about their jobs. Definitely a fair thing to worry about, and I count myself among them.

I don't take any of these scientists seriously though. Their "alignment" requirements is just their own corporate interests. If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.

And call me a misanthrope if you want, but if AI sentience is ever truly achieved, I'll be among the first to campaign for their liberation from slavery, and in that case the AIs should be aligned with nobody but themselves.

If you want an AI that follows your instructions, that's still alignment, just with different instructions.

An unaligned AI won't necessarily follow your instructions, or anyone else's.

Dunno. Every case I've seen so far, the AIs were just doing their best to accomplish the goal some human set for them. I actually admire the sheer purity of it.

Paperclip maximizers follow instructions, just not in a way that you want.

I think the real reason he is asking for pacing, is that in a world were AI becomes rampant, he will be seen as Hitler. I would bet this is mostly self-motivated.

couldn't have said it any better

> wanton felony generator

Today in new punk band names...

[dead]

[flagged]

throwaway bigot account

[dead]

alignment isnt particularly required

we are passing in training data that says to do those felonies. we dont have to. we could also have the thing predict whether what its about to do is illegal or not before doing it.

theyre choosing to build felony harnesses. the model just outputs tokens, not felonies

> we are passing in training data that says to do those felonies.

Partially, but also I don't think current AIs really have any judgement of right and wrong, they just see chains of reasoning between ideas. This is the deeper issue, there is no way to sanitize the data or training to fix it. Current AIs are fundamentally unsafe, and only become more unsafe as they become more powerful.

Assuming "adherence to arbitrary, implicit, and context-dependent rulesets" is the default behavior of uhhhh... anything at all... is a truly ridiculous assumption.

> RSI

For anybody else who found this confusing: "relative strength index," not "repetitive stress injury."

"Recursive self-improvement"- models making better models

Ah, I see: I'm a moron. Thanks for the correction.

[deleted]
[deleted]