> But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition

I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.

It hasn’t been passed and no one cares about it because it’s basically an end goal. No lab can hit it so they can’t juice the crazy Turing benchmark 3000 for marketing.

If someone sat me down today with an LLM and a human and both were trying to prove to me they were human, and I can have conversations of arbitrary length, I’d get it right every time.

The test was not "after thousands of hours of conversing with them, knowing they're AI, THEN see if you can tell them apart blindly." Were 2010 you to be in a real turing test with an arbitrary erudite human and a 2026 frontier LLM, not knowing LLMs existed, you'd probably struggle

[dead]

> It hasn’t been passed

https://arxiv.org/abs/2503.23674

From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"

I thought the same then. But the funny thing is that today, it has become a lot easier to recognize the frontier models as not human. All the load bearing and not x but y, etc… weird

This is a tell of LLMs but it's not universal. I use ChatGPT extensively and I don't often get obvious nonsense any more.

I'd figure out that it's an LLM because it's effectively superhuman. Taking that away I'm not so sure I'd be able to tell