Probably too late for this, but I have argued before that language is a fundamentally lossy encoding of the human experience. We do our best to describe what we're seeing and experiencing using language, which is fantastically expressive, but it has its limits. I think we see glimpses of this when we find ourselves saying things such as, "it's impossible to put it into words" or we overload certain words when we mean very different things, such as, "I love my children" or "I love apple pie". Clearly the word "love" here has a certain magnitude that is not being expressed, yet it is understood by the listener somehow.

So, I do sort of buy into this idea that Einstein was simulating the world and running experiments on those simulations in ways that were beyond what you could encode in natural language. Will AI be capable of doing this, if it is bounded by training data that is composed almost entirely on language? One might argue that if AI is training on a lossy encoding/representation of the human experience, how will it be able to simulate anything beyond that experience? Unless it does so in a way that we manage to do when we image objects beyond 3D. But now I'm just rambling.

The lossy compression of language is why we should find it unsurprising that LLMs tend to perform better at code, than at human language tasks or reasoning. While there can be subtle semantic differences in real codebases (using "null" to mean "unknown" in one context, versus "intentionally blank" in another), there is a much tighter coupling of semantics to meaning (low ambiguity) compared to "love" in English (let alone any inexpressible je ne sais quoi).

With apologies if this is common knowledge at this point, 3Blue1Brown has been doing an excellent series on compression, and its relationship to intelligence (or more controversially, that they are one and the same): https://www.youtube.com/watch?v=l6DKRf-fAAM

But that also throws in sharp relief, that there is vastly more to the human experience than intelligence alone: qualia, desire, gut instinct, intuition, emergent creativity. (Whether the "God of the Gaps" for the delta between capabilities of human vs AIs is fixed, or diminishing, or even shrinking to zero, remains an open experiment we're all living through.)

> Probably too late for this, but I have argued before that language is a fundamentally lossy encoding of the human experience. We do our best to describe what we're seeing and experiencing using language, which is fantastically expressive, but it has its limits.

Yes, and sometimes this is very intentional. Take for example a short poem which if you sit and really think about it for a long time, you could go off on a mental tangent of imagining what sort of kingdom or empire created a statue that is now "two vast and trunkless legs of stone", for instance. Being terse and allowing for human interpretation is kind of the entire point of something being written like this.

I met a traveller from an antique land

Who said: Two vast and trunkless legs of stone

Stand in the desert. Near them, on the sand,

Half sunk, a shattered visage lies, whose frown,

And wrinkled lip, and sneer of cold command,

Tell that its sculptor well those passions read

Which yet survive, stamped on these lifeless things,

The hand that mocked them and the heart that fed:

And on the pedestal these words appear:

"My name is Ozymandias, king of kings:

Look on my works, ye Mighty, and despair!"

Nothing beside remains. Round the decay

Of that colossal wreck, boundless and bare

The lone and level sands stretch far away.

> Being terse and allowing for human interpretation is kind of the entire point of something being written like this.

It's also to a certain extent why LLMs work.

When IBM Watson was playing Jeopardy, one of the game prompts was:

> It was the anatomical oddity of U.S. gymnast George Eyser, who won a gold medal on the parallel bars in 1904

The man was missing a leg and used a prosthetic. Watson's output was, "What is leg?"

At first it was regarded as correct. If a human said that you could conclude that they knew the answer. But then the judges decided not to give Watson the point because its output didn't provide enough specificity to prove that it understood the context.

If you ask an LLM what kinds of things taste sweet it can give you examples like cotton candy or strawberries, but it has never actually tasted anything. All it knows is that the training data contains the association between those tokens. But the human reading the output knows what strawberries are, which is what allows the output to be meaningful.

> The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be “voluntarily” reproduced and combined… The above-mentioned elements are, in my case, of visual and some of muscular type. Conventional words or other signs have to be sought for laboriously only in a secondary stage, when the mentioned associative play is sufficiently established and can be reproduced at will.

—Albert Einstein

Quoted in Using Spaced Repetition Systems to See Through a Piece of Mathematics,

https://news.ycombinator.com/item?id=18895613

which describes the author's experience that if you approach a field obsessively enough, eventually you begin to understand it at a level deeper than language.

If an LLM is big enough, I imagine something similar is happening.

Or, ask yourself the following question:

If you read every single book there is about The Grand Canyon, and watched every single video and/or documentary about The Grand Canyon, do you believe that you have fully experienced The Grand Canyon? Or do you just have to be there to fully experience it.

I dunno. Substitute in whatever you want for "The Grand Canyon". Maybe climbing Mount Everest or walking on the Moon. The point is that maybe the human experience is more vast than what is written about it.

In this context that isn't the right question.

Rather one should ask "Would you be able to answer any question and fulfill any request about the Grand Canyon given to you by others in a way that will be indistinguishable from someone who has been to the Grand Canyon"

This is setting up a contrast between long term memory and an LLM.

You might equally contrast someone who is physically in the Grand Canyon with an LLM given access to a drone with sensory attachments. I'd bet that if you told groups of humans and drone-controlling AIs to find something interesting, the humans would be more variable and cover a wider range of interesting discoveries but often just give up on the task, while the AI would be more likely to find something but would cluster around certain discoveries and make fewer overall.

What is the current ratio of opened to unopened cans of soda in the canyon, and how has that ratio developed over the last thirty days?

I'd be fantastically surprised if you could provide a precise answer to any such question without extensive security footage that doesn't exist.

An estimate, perhaps, but you probably couldn't even answer a simpler "how many people are currently in the canyon" without resorting to estimation given the lack of information currently available via the mentioned sources.

I suspect the same limitations apply to Grand Canyon visitors

I think I could get very close via video. But that depends on a lot of my non-canyon experience moving around the world, and LLMs are very flawed in their ability to input video, so I think they would not get nearly as close.

I've been to a few natural wonders (including the Grand Canyon) that I saw in advance on video. At least for me it isn't close at all. Even if audiovisual elements could be near-perfectly reproduced by video (imo not even close with modern tech, no screen is capturing the brilliance of sunlight), you aren't capturing the temperature, the feel of wind or rain, the smell of the plants around you, etc.

Well, I've never been to the Grand Canyon myself, but what you said perfectly aligns with my intuition.

Also, you touched on how even modern tech still falls short of the true experience. Just look at the history of motion picture since the early 1900s. We have added sound, color, bigger screens, faster refresh rates, more pixels, 3D (sort of), surround sound, spatial audio, IMAX, etc. Almost seems like video leaves a lot to be desired.

4d gaussian splats are going to be a major revolution in vicarious experiences

But you already know those feelings from being outside elsewhere. You get a unique combination there, but that's only worth so much.

> only worth so much.

Worth... quite a lot imo.

First, in the slightly objective sense of "What is this place like in real life, under X conditions." But, more subjectively, watching a video of a glacier and imagining the wind/rain/cold doesn't even approach 5% of the intensity and awe of climbing up a mountain yourself to see a seemingly endless expanse of ice, struggling to stand steady because of the wind, shivering because of the cold and rain. And finally, in the financial sense, a lot of people routinely spend many thousands of dollars and days-weeks of their time to experience natural wonders in real life.

There's so much natural wonder in the world in and around your home.

It's not that the value of the grand canyon is low in an absolute sense, it's that in a relative sense the gap between "empty numb void" and "hiking at home (plus grand canyon videos)" is far far greater than the gap between "hiking at home (plus grand canyon videos)" and "actual grand canyon".

> in a relative sense the gap between "empty numb void" and "hiking at home (plus grand canyon videos)" is far far greater than the gap between "hiking at home (plus grand canyon videos)" and "actual grand canyon".

Have you ever been to a place which is really high up? You're above the trees and can not only see as far as the curvature of the earth allows you to, but are high enough that the curvature of the earth allows you to see for miles. The air is thinner so it's easier to take a breath yet harder to catch it. Your sweat evaporates faster but a breeze doesn't cool you as much. The sun is brighter, rises earlier and sets later. The wildlife makes unfamiliar sounds and sound itself is changed. Even the same food tastes different.

It's not something you can get by sitting at home and changing the background on your computer.

The experience of every famous place I've visited after watching extensive video before hand, has been very different from what I imagined from the video.

No way. You get to the rim and only then you realize how huge it is. And then you still have to go down into the canyon.

The experience of vastness is mostly due to (high-res) stereoscopic vision. When you experience 6DOF VR in a fairly high-res world it definitely comes much, much closer to the awe one feels with the real experience.

It's similar in dismissing artificial audio based on listening to some pop song on a basic mobile speaker versus listening to well-designed binaural audio on good headphones. The latter can sound 'freakily' 'real'.

I think this is very critical concept and existing gap which does not makes into AI conversations. Interestingly enough a lot of Si Fi movies have captured this where the AI starts to feel and have "thoughts". Have to give kudos to the authors for being so creative and imaginative.

bro, it is a token predictor. You had me in the first half, language is humanity's greatest invention and all of the LLM "wins" can be attributed to existing language(imo). But no it's not going to have a vision like Einstein because it doesn't have a brain it is a token predictor.

Yes, and a human is a procreation machine. Bro.

[dead]