I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.

> If this is truly AGI (subject to one's definition of AGI still)

Scoring well in a benchmark that's called AGI does not make an LLM AGI.

The goalposts of AGI will shift forever. If you showed our current capabilities to someone from 2016 it would be declared AGI.

Is anyone from 2016 still alive today?

If so I'm hoping we can track them down and have them tell us if they think this is AGI.

I'm still here, nearly 50 years and counting. If you had asked me what I imagined AGI would look like back in the 90's, I would have told you "A system that can do everything we can: see, hear, think, do.". If you had shown me GPT-6 back then, I would have said "It looks like a really powerful program, but that's not really what I had in mind.". That's AI, but it's not quite general.

And then you'd ask it about an area you're knowledgeable in and realise it routinely makes stupid mistakes.

Or you'd ask it to add a new page to your website and shout at it to use your existing brand colours instead of inventing some and realise it's not AGI at all...

As someone who spent countless nights tweaking Edge Detectors (looking at you, Canny), morphology operators, etc., building models to recognize 10 handwritten digits, let me tell you: the current set of LLMs (even the smaller ones) seem like magic. I had never imagined a computer would do such things in my lifetime.

Exactly people can say whatever they want, but current level of LLM is AGI level to me. It is already on par with senior programmer if the instruction/prompt is right.

Once we have 1000 tps, i am sure robots etc.. will also start working like magic.

I don’t know, I am writing a modest 30 page paper with Fable and even after rounds and rounds of feedback and improvements there are so many things that are just plain wrong or weirdly out of place or just stupidly written that Fable 5.1 doesn’t seem to have any awareness of by itself that I don’t think it’s AGI, I think a human researcher can easily outclass it in writing and problem understanding. It definitely has super human capabilities but it lacks awareness or self reflection in my opinion.

For example it should be easy to tell it to not write a paper in the style of a clickbait SEO article or use all of its stupid hallmark AI writing patterns “it’s A, not B!” And a smart human that would be told that would be easily able to comply with that but the model needs to be told in a very detailed way and it seems to lack even basic capabilities to reflect on this, when explicitly given a sentence it will be able to rewrite it but otherwise it’s mostly blind to it. That’s to me a hallmark of it being overtrained on the specific tasks or problems so it appears very smart but once you go off script it still shows that it’s not a “real” mind.

Of course it’s amazing and has super human capabilities in many areas but if you honestly think it’s better than Einstein like some people suggest why can’t it write a simple “good” academic paper even after giving it specific examples and instructions.

Maybe that’s what makes these things dangerous, they have super human capabilities in some areas but apparently lack self awareness, taste and meta reflection abilities. The only reason people aren’t afraid more is that they don’t act in the physical world yet, imagine giving it a body, superhuman strength and letting it care for your child when it has a strong “urge” to comply with your exact request and little to no self awareness and human basic instincts.

100%

"You're absolutely right to call me out on that. I shouldn't have stopped the baby crying by killing it, that's on me."

Doesn’t the G stand for “general”?

An AI model that’s human-level at programming is an incredible achievement. But it isn’t general intelligence. It’s highly specified intelligence.

>It is already on par with senior programmer if the instruction/prompt is right.

It's magical to me as well, but I don't feel like it's AGI.

Because in my experience a Senior Programmer does not need the right prompts to deliver the right outcome! :-)

Compared to what we had in 2016 with RNNs, this is effectively “AGI”

OK, so its way better. that doesn't make it AGI.

If I can't give it an arbitrary task and have it solve that task eventually, it's not a general intelligence.

Are you guaranteed to solve an arbitrary task eventually?

I believe so. AIs are shockingly good at a lot of domains, but there's still a lot of pretty basic stuff they don't really "understand" at a conceptual level and (currently) they can't learn to get better at them.

(obviously it might take years for me to get good enough at something, or if you set the "arbitrary" task as something ridiculous, but lets work in good faith here and think of something the average human could do after learning about it)

If we progress to the point where an LLM instance can meaningfully learn to get better at something overtime without retraining, then I will accept that is basically AGI. Right now, they still seem to be pretty boxed into their training, even if you can prompt them to act differently.

True. "AGI" has also become a marketing term. Achieving AGI has become valuable, so companies will move the AGI goalposts, over and over again, so they can achieve AGI, over and over again.

If you came at it from the perspective of imitating what the human brain does, we now have a very very powerful speech center and short term memory, and vision catching up. The other parts are missing. I‘m sure that’s being heavily researched.

I hear the T-rexes were still roaming the earth trying to eat us cavemen in 2016.

If i suddenly travel to 1500s i would also be considered genius(in some way)

I bet you’d think they are not even conscious.

talking about self proclaimed, it's about as much AGI as openAI is open.

But they declared it...

“I declare bankruptcy!” - Michael Scott

I DECLARE AGI!

"Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"

[deleted]

did they?

"""

In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”

"""

What test do you propose as the actual go/no-go gauge to verify if some model is or is not AGI?

If you’re talking some nonsense, silly singularity… than whatever, don’t care.

But if you’re asking when a model has a sustainable general intelligence, for me, it’s pretty easy…

When it makes financial sense to run it 24 hours a day.

I can have Astra run a large-scale infrastructure migration 24/7 (much of the time waiting for results), completing it in weeks or even months faster than I could before agentic AI.

What does it mean to run a model 24 hours a day?

Aren't we way way past that already? QPS to any of the frontier models for a given point in time is most likely (far) greater than zero.

For whom? That is a fantastically ill-defined test. Everyone here is comfortable throwing around this or that is or isn't AGI which is fun because, at the same time, nobody seems to have a testable definition.

It makes either position pointless to argue.

I mean the laundromat runs the machines pretty much 24 hours a day but a washing machine is not AGI.

Directly - something can be useful without being AGI.

[deleted]

There can be no such test because “AGI” is (or has become) a pseudo-philosophical/socio-political concept rather than a scientific one.

It always has been. The idea around here that we can actually define intelligence and point to it is sophomoric and incredibly frustrating.

If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.

Populace

Hey now! Keep your reason out of their marketin^H^H lies!

We've had AGI (artificial general intelligence) probably since the first release of ChatGPT, and certainly since the first agentic harnesses. They're just finally acknowledging what the term means.

We've had AGI since RNG! Cut the poor, unacknowledged RNG AGI some slack, will ya? It can literally solve everything when you're patient enough.

There's so much that the term includes that isn't even feasible with an LLM

Artificial. General. Intelligence. The ability to solve (even partially or even badly solve) problems drawn from arbitrary problem domains without pretraining on the specific problem class. You can pose any problem of any type using natural language to an LLM and it will attempt a solution. That's literally all the term means.

You (and the rest of the media and many industry figures) are conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).

There's a simple rule: someone of higher intelligence can successfully pretend to be someone of lower intelligence, but not the other way around.

Humans can successfully pretend to be LLMs, but LLMs are still not successful at pretending to be Humans.

ASI "matches or exceeds humans in at least one field"

AGI is defined as "matches or exceeds humans in all fields"

Superintelligence is defined as "strictly exceeds humans in all fields"

> conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).

So now humans is "super" intelligence? it's nice to move the upper bar so that more stuff can be called "just" intelligence.

Reference class in this case means not an example but what the comparison is against. Superhuman means better than humans. General intelligence is defined without any reference to human capability levels.

How can somebody define intelligence without a reference to human capability? Humans are the one judging it.

general intelligence for beavers or a birch forest would be very different than general intelligence for humans...

Intelligence is problem solving. It can be defined in terms of optimization theory.

Which is also a uniquely human take...

I don't even know what you are arguing for.

I don't think we'll be able to meet, and that's ok since the definition isn't universal. I align more with Demis Hasabis' views on this

Very valid point, shame it’s buried so deep in the comments’ tree.

This is a very mundane release compared to GPT-4 and GPT-5. I think they probably scaled back a bit after the lukewarm response to the GPT-5 announcement. But it still very weird that there wasn't even a livestream,

There is simply no level of announcement that won’t have people complaining. What is so important of having a livestream?

Seriously, if they’d done a huge splashy launch, we’d be reading one hackneyed comment after another about their fake hype or whatever.

I think we're getting to the point where it is difficult to identify the goal post of AGI.

Is it rapid skill acquisition? -> ARC benchmarks are saturated Is it breadth of knowledge? -> See many ... many benchmarks Is it ability to do hard tasks? -> see terminal-bench and released outputs.

We are at the point where the starting point for most tasks should be "send your agent to work on it."

So where do we draw the line in a way that doesn't move every 6 months?

The real answer is converting from any format to any other reliably. Text to speech, speech to text, music to video, image to 3D, piloting a drone by converting video feed to rotor speeds, literally any file conversion, like html to pdf, photoshop project to png, png to photoshop project,... turning Toy Story 1 into a series of Blender scenes with all textures, models, materials, lighting, camera movements matched to a tee, should solely be a matter of how long you let the model run. It should never run itself into a dead end. It should instantly know when it is making mistakes, with no human babysitting it.

I can do none of those things.. I hope that I am generally intelligent.

1 year ago we viewed models as tools and agents were just kinda toying around, that we now think the bar is literally an anything to anything converter through one agent is wild.

It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.

MinMax

• 98.6% on ARC-AGI-3

• 97.6% on frontier math

• 95.9% on CAD

• 100% on ExploitBench

Nothing modest about it

Except the release announcement. You know, the thing the OP you're responding to is specifically pointing out?

If a video announcement and a press release would change a person's mind on whether this is AGI, I don't put a huge amount of weight on that person's conception of what AGI is.

There has stopped being a formal procedural consequence for OpenAI leaders to declaring AGI, there is a clear (small) business benefit to doing so, and the capabilities of all the frontier models are impressive. So why not declare AGI? It's not like anyone can prove it's not...

Don't be surprised to see other (or even the same) people declaring AGI again and again, as it becomes the best time to do so for different parties.

> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.

Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc

I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

I don't remember where I heard this, but one of my favorite criticisms of the current AI situation is that it's wrong simply because of the size and energy required compared to the human brain. The idea is that there's still some element missing thats fundamental, and that the way we train them now is part of the solution, but not all of it. I think finding the extra missing element is going to take an entirely different approach that will also solve the sizing and resource issue. The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.

Yes, the very explicit plan of both OpenAI and Anthropic is to use the not particularly efficient LLMs to automate their own AI engineering. That seems to be going well - on coding front and model tuning front so far. They have more planned.

And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain.

It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it.

It's scary, TBH.

I recall them saying they use models to write CUDA kernels and whatnot. Makes sense, and unsurprising that models are good at writing code.

But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.

Quite the gamble.

> but I didn't think it'd solve maths problems

Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.

It’s a bird! It’s a plane! It’s…AI skeptics moving the goalposts at light speed!!

What about finding the 1st known complex structure over S^6, proving Ehrhart’s volume conjecture, proving a sharp "density" bound on primitive sets conjectured by Erdos >60 years ago?

> Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.

It's not good to be confidently wrong the way you're being.

We've had mathematical problems solved by brute force in the past.

We've then improved that through systems similar to prolog intentionally searching a tree.

Then systems added heuristics for which paths in that tree are likely to be taken.

The LLMs are just using slightly more accurate heuristics for this task.

But the real measure of understanding are tasks that are not so strictly constrained.

I mean, the plan is to use these models to find and solve those gaps. That's kind of the whole pitch of these companies: they spend a TON of money upfront setting up this infrastructure, but each iteration yields a system capable of making the next iteration even better.

>The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.

Perhaps. But only at that point, not leading up to that point.

It's kind of like setting up scaffolding to build something. You spend all of that time and money to build something just to tear it down in the end. But the point is that it's simply a cost to be able to build the actual thing you're building.

If these companies are able to achieve the results they're looking for, none of the investors involved are going to care that the datacenters and infrastructure they spent so much money.

If we manage to get to AGI and it looks, works and behaves like a human brain... I mean, cool, but that's a very useless AGI compared to the incredible stuff we have access to today.

The HN crowd I'm sure will still be unhappy calling it AGI because "it's not AGI unless its speech comes from the cerebral cortex region of the brain, otherwise it's just sparkling emoji" or something.

I think the idea is that you wouldn’t need humans to do anything anymore, right? As impressive as it is, it’s still ultimately directed by human planning and coordination. Assuming they are aligned, you could have a collection of AGI that you let loose and they tirelessly solve all of humanity’s problems, do all of our work, and progress science and our understanding of the universe.

Those are all things that humanity is doing everyday. What we have is amazing, but it’s not that.

The infrastructure you just described is absolutely buildable today if you’re willing to burn tokens.

you're on one of the most pro AI spaces on the whole internet and yet you're still crying about "the hn crowd", what a bizarre distorted perspective

>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.

We have not "hit a wall" by any stretch yet. I don't understand how someone can even hold this viewpoint? It's mind boggling.

Harnesses magnify and make the intelligence actionable, but we have not reached limits on raw intelligence yet, not even close.

Agreed. Trivially observable by using a frontier model from today and one from 6 months ago with the same harness.

I don't think so.

One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.

I think models using these harnesses were also RLHF'd hard on responding to looping instructions and following through on goals. Older models were tuned for basic chat responses.

If someone could tune models of that size to have comparable effectiveness at much much lower costs, they would have done so by now.

"The harness improvements are the real sauce" is like a sincere "It's gotta be the shoes" take about Micheal Jordan.

(For the younger: that line was from a series of Nike ads where his skills were being explained)

We call that a sigmoid.

People really believe in this AGI marketing?

> No video announcement

They've released two videos:

Vision video:

https://www.youtube.com/watch?v=1QNsdr-Qx_I

(kinda reminds me of these retro videos about the future home: https://www.youtube.com/watch?v=rnbaehgxdp0) ((can't find the other one where someone controls the home computer with voice))

Vibe coding with it:

https://www.youtube.com/watch?v=-TTyyY3VWh8

Given the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to continue to increase their market share from Anthropic.

They’re really, really scared because of the Mythos controversy. Skynet will be under hyped.