Here's another definition of AGI from Sam Altman:

https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...

Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?

Kevin Roose (New York Times): I probably would, yeah. Would you?

Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.

So, "really good" is the boundary now, whatever it means.

What's your definition of competence boundary for human coworker?

If the new AGI benchmark is "be Einstein/Feynman" then we've hit AGI.

What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?

The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.

> What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes?

Then it’s an expert system.

Stephen Hawking wasn’t very good at folding clothes.

The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?

You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.

Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows.

I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech.

Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence.

The only important part of 'general' is the ability to learn from experiential data and update your own model. That's what leads to general capability. Humans can't oneshot any task natively, but we can practice for a while until we uncover often novel methods of accomplishing something.

Therefore: the current transformer architecture is fundamentally incapable of AGI because the models have no mutable long-term memory.

You only have weights (large immutable memory), or context (small mutable memory).

Humans have mutable long-term memory: I can learn a new skill, adapt an old skill to new information, or learn new knowledge today that I couldn't perform/didn't know yesterday. I don't have a training cutoff.

Context engineering is an attempt to paper over this limitation. You can get really far with context engineering and huge models, but you will never get to AGI because there are many tasks where humans' mutable long-term memory outperforms.

For example, a human can invent a new musical instrument and then learn how to play the instrument they just invented. That's inference (inventing an instrument) leading to training (neuroplasticity). Humans have the ability to train our NNs with considerably fewer training samples. Everything that you can do with transformers is in one causal direction: training -> inference.

So if we take a huge with enough compute (CPUs, b200s, petabytes of SSDs), we install on it both the Astra, and the toolsuite to incorporate new sensory inputs (threads/sessions), camera, microphone, temp sensors, the lot, into a new version of the model. This model is then swapped for the old model, or traffic slowly brought over, or even adjusting weights in place.

Then my hypothesis is that thing as a whole could achieve AGI.

This feels like a very close approximation on how we humans evolve our brain. By encountering new experiences/sensations, classifying them as negative or positive to us, filling it away in neurons. Or by training motor skills etc. In the end we get more connections between neurons in our brain and we are capable of more.

Bingo, LLM architecture just does not lend itself to becoming AGI. They can get really good, sure, but they will always struggle with novel input and scenarios.

The more training data that is shoved in to them, the more they'll seem to solve novel situations, but in reality it'll be things that exist in the training data.

[deleted]

Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation).

Most people can't draw a bicycle. There was an artist 10 years ago that asked people to sketch a bike, and then turned these sketches into 3D renders - quite funny.

https://mymodernmet.com/gianluca-gimini-velocipedia-bicycles...

https://qz.com/681345/an-artists-3d-renderings-of-bicycles-d...

How good was Einstein at drawing pelicans on bicycles by writing SVG code?

Checkmate, meatbags.

Pelicans are a solved problem at this point. An open-weight model on my own machine gave me this: https://crimson-jeri-74.tiiny.site/

And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans.

Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed.

> Pelicans are a solved problem at this point. An open-weight model on my own machine gave me this

That feels kinda like when I remember seeing Ocarina of Time for the first time, and thinking “oh my god, this looks just like real life…”.

For me it was the wheels. I couldn't stop staring at the wheels... how did it get them so freaking perfect? Mad respect to GLM 5.3.

Laundry folding has become a doable demo for startups, and ChatGPT has been spitting out college essays for years.

I can't draw a pelican. Literally my only point of reference would be AI pelican drawings from the test. Otherwise I wouldn't know how to draw one at all.

I would be able to draw an accurate bicycle, but I'm an outlier on that. Most people could not draw one [1].

[1]: https://www.booooooom.com/2016/05/09/bicycles-built-based-on...

Call me when Astra gets the Nobel prize... We'll have AGI when prizes have two categories, one for assisted humans, and one for pure-AI.

Wouldn’t that mean producing novel work like relativity and QED?

I would maybe argue that Einstein was the most LLM-like of great thinkers.

A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.

A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.

That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.

[dead]

What is novel physics?

I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own?

This is as good a time as any to note that we might be closing in on a new conceptual revolution in our own time as it relates to holography and an information centric approach to spacetime. Obviously It's the furthest possible thing from a guarantee, but it has much of the enthusiasm and motivation that string theory had previously enjoyed in prior decades.

So it could be a natural experiment for whether AI can contribute to novel physics. Specifically, there's a big question about weather. Something like our informational understanding of black holes where information inside it is equivalent to information on its boundary (which I'm sure I'm not saying correctly), might be generalized to regular space-time. More people should be freaking out with excitement about this and perhaps it's something to which AI can contribute.

Do you have anywhere you recommend where I can read more on this?

Honestly I don't think I have single great article, though some Quanta ones are ok, and the Wikipedia article is okay.

The best thing I can recommend is what I did, which is ask Claude about the significance of (1) quantum computing error correction, and (2) error correction in black hole holography and research convergence between the two.

https://en.wikipedia.org/wiki/Holographic_principle

https://www.quantamagazine.org/how-space-and-time-could-be-a...

Edit: this whole article, despite it's boring title and hook, is maybe the best discussion of holography as a recent and active research frontier.

https://www.quantamagazine.org/if-the-universe-is-a-hologram...

There was symbolic AI programs in the 1980’s that “discovered” Kepler’s laws and the resulting solar system model from just tycho brache’s astronomical observations. That was the the very first “new physics” ever.

i wonder if we could train a modal, and omit all data prior to 1899, and see what happens?

Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this.

as always, as good as its prompt...

I assume solving one of the major open problems of physics?

Would this be possible without it being able to run novel real-world physics experiments autonomously?

(Note: I am not suggesting we let it do this. Please don't, in fact)

AI could discover candidate novel physics without autonomously operating new physical experiments, and humans or instruments can later independently validate the result. This is analogous to how Einstein developed theories whose predictions were confirmed by experiments and observations only years or decades later.

Do we have any examples of an current day AI system introducing a novel concept or perspective. We've got plenty of counterexamples discovered and some theorems proven, but afaik nothing analogous to a new definition.

Pharmaceuticals already are…

It could be a good theoretical physicist. Actually it could be a good experimental physicist as well since senior experimental physicists use grad students for the manual labor.

Why not?

That's how we get AM.

That’s already been done. I know of at least one novel result contributed by Claude to frontier physics. I’m sure there is more.

For it to be like a human it wouldn't just need to solve existing phsyics problems, it would need to push the field forward and introduce new paradigms.

Solving "open problems" will push the field forward.

My comment wasn't very long, yet you somehow still ignored the main part, "and introduce new paradigms". The point is whether it can do everything humans can, entirely new theoretical frameworks and ideas, such as string theory or dark matter, are not coming out of AI at the moment.

Creating new physics is the new AGI goal post