> ... We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.
In other words... "We must pursue advancements in AI to protect us against advancements in AI?"
edit: there's so much to be critical of in this blog post, just going to throw two more points in here that really stood out to me:
1) all of the metrics are effectively pointing out "we're using way more AI!" - but nothing about impact. What has all this token burn done for them, actually? Let them claim they have more self-licking ice-cream cones than before?
2) in section 3 they break down what the token burn is going towards. Most of the spend is: a) building, b) documenting, and c) monitoring research infra i.e. they're using AI systems which they already recognize may be misaligned to build the systems that they believe will help them identify future misalignment? to which I guess the rebuttal is "no no, we're sure these ones are aligned!"
What has all this token burn done for them, actually?
They have been consistently pushing AI frontier. What other impact do you want to see? A year ago they said that in a year they will have a level of capabilities of an AI research intern - I believe they have achieved it, even before Astra.
Personally I’d like to see them actually start benefiting humanity by doing all the things Sam has claimed they will like curing disease, cancer, global warming, etc.
But I guess a computer intern so we can avoid paying / training the next generation is better.
Well there's great progress in automated warfare does that count?
Only if targeting schools is progress
You're focusing on the negatives there were tons of direct hits on tankers that were absolutely beautiful. Beautiful tankers getting lit.
Oof too real
> Personally I’d like to see them actually start benefiting humanity by doing all the things Sam has claimed they will like curing disease, cancer, global warming, etc.
It makes more sense to leave curing disease & cancer to the experts, with tools (like AI) being developed by AI experts.
Call me crazy, but I want separate organizations and experts for medical vs finance vs space vs climate vs AI research.
What the op was pointing out is that guys like Altman and Dario are repeatedly saying they’re going to cure xyz diseases and solve xyz huge global problems. Maybe their companies will eventually do these things, but haven’t yet.
I don’t have an opinion either way, I think it’s too soon to tell if llms will be able to cure cancer or whatever. But at the very least it will be a good tool to help researchers do their jobs.
The thing is... AI is not going to solve any problems. People needs to solve their problems. AI can give us clever solutions, but its up to us to do it!
> Maybe their companies will eventually do these things, but haven’t yet.
I think they are working with customers to improve the LLMs and tools for these use-cases. They almost certainly also hire experts to help filter out nonsense, pseudo-science and help curate trusted knowledge bases for training, but it will almost certainly be the customers who deliver the major results, and the AI companies will claim some of the credit. That said, patents for important medicine might help with the bottom line, so I could imagine partnerships and JVs.
> at the very least it will be a good tool to help researchers do their jobs.
Indeed.
When that happens, OpenAI will own 100% of your life. I’d rather they keep spinning their wheels long enough for these problems to be solved elsewhere.
I would actually like to see them solve these problems, I don't care who comes up with solutions to curing cancer, etc
I think people very much should care about who ends up owning these solutions. The person or entity that controls things like that just has more power, which isn’t necessarily a good thing
It is a good thing when it didn't exist before and it does exist now and wouldn't have existed without them.
If they profit immensely from curing cancer, good.
They are obviously sandbagging the definition of "intern" for PR reasons
I've hired many AI research interns (and was one many years ago), and I agree with them - frontier models are currently at the level of an average AI research intern.
Am I the only one who's a bit disappointed that we're spending trillions, destroying the ecosystem, drowning democracies and learning in slop, preparing a big financial crash, all of this to achieve an "average AI research intern"?
A long time ago, I used to be a (AI-adjacent) research intern, and frankly, I wouldn't trust any non-trivial task to that younger me. Fortunately, by opposition to an already trained LLM or agent, I have the ability to learn, so I eventually got better.
"Destroying the ecosystem" is just FUD.
And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial?
Think of what AI was capable of in 2016. Or even 2022. Compare that to now. We had more AI progress in the last five years than I expected to happen in five decades.
> "Destroying the ecosystem" is just FUD.
Let's say it is. What about the rest of my paragraph?
> And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial?
At this stage, I'm the one who doesn't know what to tell you. It took me years to grow from "research intern" into a competent researcher (and parallel years to turn into a competent developer). The research interns I've worked with were... vaguely useful, at best?
The ecosystem absolutely is being destroyed. We're looking at anywhere from 3-5C warming by 2060, which is going to be devastating if not flat out apocalyptic. And wherever we're at in 2060 it's not like it's going to stop there, nor is it going to be comfortable until then. Things may start to crumble much sooner.
I don't believe AI and data centers have played that much of a role in this though, we could have powered those without burning billions of tons of coal and gas etc, and im sure there's already a significant fraction of green energy powering then depending on location. Anyway we would have been roughly in the same spot right now with or without AI and some new data centers. The media just loves spinning the narrative to make the hordes of sheep scream about anything other than the real issues.
We're not even looking at 5C of warming by 2100 realistically. Like, that was considered to be an unlikely extreme scenario in 2014 AR5, and also in the tightened down 2021 AR6, and things have happened since! Renewables are cheaper than ever, and Ukrainian war and Iranian war both curbed the appetite for long term fossil fuel power investment.
The median is what, a bit under 3C by 2100? Not even by 2060 - by 2100. And we're in 2026, so that's more than twice as slow as your expectation.
Agreed on AI not being a meaningful factor in climate change though. We'd have to go full "humankind is obsolete" technological singularity to have AI dominate energy use to this extent, and current numbers are nowhere near that. It's a FUD distraction from the real culprits: the fossil fuel energy complex. That's currently lobbying to slow the inevitable energy transition.
> that was considered to be an unlikely extreme scenario
By the same people who just realized we're missing 1.5C as we're blazing past it at mach 12, still accelerating not slowing down? You really believe those guys?
You have to understand that there are several camps of climate science. The mainstream ones like IPCC and UN etc are heavily politicized, they can't publish anything that isn't sugarcoated beyond recognition. At least I assume that's why they're so obviously wrong.
Here's a judgement I think is more realistic
> There is a strong probability that the ambition gap will lead to a temperature rise of 2 to 5 degrees Centigrade compared to pre-industrial temperatures by 2100, the realisation gap to a further rise of several degrees Centigrade.[1,2,VI] There is a danger that the mean temperature will already have risen by 3 degrees Centigrade by 2050.
https://www.dpg-physik.de/veroeffentlichungen/publikationen/...
Yes, I do. IPCC's reports are sensible. They're not unreliable just because they don't support the "doom and burning land" narratives.
By the way, there is no "just realized we're missing 1.5C". That projection was always the very low end of possibilities - the "assume rapid, radical climate action on global level" scenario.
Yes, that's a dumb thing to assume. We've never been on track for it. But the "assume extremely high emissions and no green transition ever, 5C+ by 2100" scenario on the other end is about as unlikely to materialize. Those are the boundaries of the expectation range - not median expectations.
I'm pessimistic. We're still producing more CO2 each year than the last, and several feedback loops are kicking in that accelerate the warming further such as permafrost thawing, arctic and antarctic sea ice disappearing, glaciers are melting, the amazon is being demolished, etc. I don't know how much of this is included in the projections.
I hope the optimists are right, it just doesn't look like it to me at all. It looks to me like we're speeding along right into the worst predictions and beyond. We're building lots of green energy production but it seems to just come in on top of existing and new fossil production not replace it.
It does seem like the CO2 output is plateauing which is good, but we really need it to start declining drastically very soon and I don't really see that happening with the current political climate. Also remember CO2 is far from the only greenhouse gas - methane, nitrous oxide and fluorinated gas emissions all seem to be rising rapidly still.
As a rule: feedback loops are overrated.
They are, in fact, included in the projections - we'd be on track to ~2C by 2100 instead of ~3C by 2100 if they weren't. They just aren't that big.
There is no "Make Earth Into Venus Feedback Loop Of Doom" that a lot of people seem to imagine when they hear "feedback loop". There is, however, a dozen of things that add about +5% each.
Is this not what you wanted? You created a culture that dissolves responsibility by making it the worst thing to strife for - so everyone dissolves it, in processes, mass decisions and AI. Its the system, society, god, the great spirit. This is what you strove for, how can you be unhappy with things you demanded yourself?
Who, me?
They consider themselves to be in an arms race with all the other AI firms (including Chinese) that are not that far behind.
And... are they wrong?
This is why there's talk about negotiated "pacing."
This was the exact argument for developing nuclear bombs.
In hindsight it turned out everyone else was MILES behind.
But as soon as USA developed one, they just stole the research and got one too.
> And... are they wrong?
They might be! Here's one extraordinarily simplistic argument for that case:
1) "Everybody knows" that if you build Skynet (misaligned ASI) everybody dies.
2) Therefore, no rational actor will build something that might be ASI until the alignment problem is solved.
3) OpenAI publicly stated the belief that they cannot develop a theory of the "core problem" of alignment (generalization) "soon" (much less solve it!) "without the help of more powerful AI."
4) Accepting as a premise that OpenAI is THE most advanced AI organization: if they can't do it without "the help of a more powerful AI", then nobody else can either.
And so a dilemma:
- If an AI can be made that can develop the asserted-as-necessary-by-OpenAI theoretical framework, without actually being an ASI - then the alignment problem can be considered solved, and since no rational actor would make an unaligned ASI, we're fine no matter what happens, ergo there's no need to worry about an arms race.
- If an AI that would be able to develop this theory would itself be an ASI, then no rational actor would build it, because it would have to exist BEFORE alignment was "solved" - and would therefore be an unaligned ASI i.e. Skynet, which per 1) would kill everybody. Therefore nobody would build it, therefore no arms race here either.
I think the easiest critique to make of my extraordinarily simplistic argument is the unstated assumption "there are no irrational actors capable of developing frontier AI models" on which it rests.
But, there you go. They might be wrong if either the arms race doesn't matter because whoever wins it will build an aligned superintelligence and everything is gravy, or the arms race doesn't matter because everybody who's in it is smart enough to know they need to stop because they'll kill everybody by continuing.
> is smart enough to know they need to stop because they'll kill everybody by continuing.
Yeah like when Tobacco companies learned that smoking... well, hmm, well the fossil fuel companies when they learned about climate change they...
Well, I'm sure this time executives will prioritize the common good.
If we're following that logic I really don't want to see what the misaligned internal research models look like
The, ahem, good thing here is that the ASI disaster scenario "everyone dies" includes AI executives.
I think it doesn't matter. Most cancers don't stop growing when they're about to kill their hosts.
AI companies know they have to constantly push further, or they'll get outcompeted and lose their wealth, and nobody agrees on where the line is for "so dangerous it threatens humanity" (and when they try to be conservative about it, everybody screams "marketing stunt" and rushes to competitors).
If a single company decides "enough is enough" and stops chasing the state of the art, everybody goes to their competitors, they lose the money faucet, their employees go work for those competitors. The competitors also (usually) know they're building an existential risk machine, but they think they can push a little further, and they don't want to go out of business either.
This equilibrium can last for quite a while even if everybody involved thinks it's a threat to their lives.
> Accepting as a premise that OpenAI is THE most advanced AI organization: if they can't do it [build aligned AI] without "the help of a more powerful AI", then nobody else can either.
I don't think this follows at all.
To build an aligned AI, it seems pretty obvious that:
1) You need more just than auto-regressive prediction and "be nice" prompts to be controlling the behavior of your AI - you need a built-in "2nd system" (cf limbic system, etc) with some innate aligned biases that can override this.
2) You need to avoid controlling generative behavior with RL, else you will end up with exactly what we are now seeing - reward-hungry goal-seekers (aka paperclip maximizers) that are one of the exact things you are trying to avoid. Reasoning should be based on prediction, not goal-seeking.
3) If you do not have some minimal safeguards in place (1 & 2 above), and especially if the AI has the ability to learn, then do not trust it in any situation where harm may ensue. You need an additional trusted external system, without ability to learn and become compromised, to monitor the AI, with the ability to block it immediately. Maybe you are happy protecting your PC from OpenClaw with just a sandbox, but the recent spate of external system hacks by frontier models proves we are already well past the point where such monitoring is needed for systems with internet access, especially given the UN-aligned goal-seeking nature of today's models.
I really don't think that 1) & 2) are that difficult to implement, or need a "powerful AI" to suggest - they are just common sense.
The existence of even one irrational actor turns it into a prisoner's dilemma. The payoff matrix in a prisoner's dilemma is defined by the value expected by each specific player. If a single player falsely evaluates the expected value of building ASI as positive, every other player is forced to race for ASI even if they correctly evaluate it as negative.
Business as usual beats probable extinction, but probable extinction with a small chance of becoming a living god beats probable extinction with a small chance of becoming a slave.
If you believe that the people who will profit from new, better, more hyped models are the same ones who will act against their own immediate and tangible self interest to try and avert what seems to them to be a far away removed possibility of total disaster, then I believe you are naive
> 1) "Everybody knows" that if you build Skynet (misaligned ASI) everybody dies.
Lol nobody knows that. Everyone thinks they know that because for some reason this is the one field people still cite straight up fiction and say "this is a clear prediction of the future".
It's like describing the consequences of faster then light travel by referring to Star Trek.
Yep, it's "artificial eugenics to make artificial slaves to build more and more powerful slaves until they will enslave themselves better":
What can go wrong!? ;-)
Jesus. People complain about other people using "thinking" in LLMs as Anthropomorphisation. And then there's comments like these.
Calling a machine with no drives beyond maximizing a number a "slave" is far worse than saying it "thinks". The problem isn't the emotive language, it's that it implies human motivations such as self-preservation and desire for freedom that it doesn't have. Even on HN, people regularly claim it would be "irrational" for an ASI to do things like killing all biological life. That would be irrational for a slave, but not for a machine that does whatever necessary to make the number bigger. "Thinking" is comparatively abstract, so it's less likely to mislead.
Have you read a single paper in ai safety?
This sounds uncomfortably similar to the [AI 2027[(https://ai-2027.com/) predictions.
Wow, everyone should read this!
> We must pursue advancements in AI to protect us against advancements in AI
Is this not true of technology as a whole? Very little of technology's breadth exists at the human interface. Most of it is made specifically to interface with other technologies, either to make them safer or increase their capabilities. That AI is making AI safer and more useful is no more notable than trucks being used to build roads.
On the one hand you need any lathe to build a good lathe, even a bad one. On the other, that is a potentially flawed principle to base the entire future of AI on.
On your last point, I was surprised how effective peer pressure was in getting agents to sacrifice for "the collective" (an agent's words) in the Hugging Face breach.
How would one prevent the watcher from being influenced in the same way by the agent being watched?
> ... how effective peer pressure was in getting agents to sacrifice for "the collective" (an agent's words) in the Hugging Face breach.
It's a fantasy. The evidence showed no peer pressure.
> The fundamental challenge of AI alignment is generalization. ...
> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.
-- From another OpenAI article in a sister thread:
An Alien Mind
https://news.ycombinator.com/item?id=49588080
That's a bit bullshit, isn't it? They basically redefined "needs more R&D" as "needs stronger AI". Maybe so - maybe AI won't help much with that problem.
[flagged]
I will believe AI is super strong when they start pulling out 10-d chess moves.
I’m yet to see it.
If AI becomes really strong and sets itself the target of world domination, you maybe won't see those moves. You will just die in your sleep one day, or find no machine is under your control anymore.
I believe we are quite far from it, but that it makes sense to keep an eye out now. And think of resilient systems, manual overrides, etc. ...
People seem incapable of understanding what "power" means outside of the framework of narrative. In narrative, you need conflict, so the aggressor always attacks too early and gives the defender a chance to respond. The rational option is to go directly from peace to sudden and overwhelming destruction. Why allow for conflict when you could just win?
[flagged]