Indeed. LLMs resemble human intelligence in more or less the same way that the output of the TI-99/4A speech synthesizer resembles a human voice.

Find memories of the prank phone call marathons during the summer of 85 when my friend got the speech synthesizer for his TI-99.

Not if an LLM over chat can fool most people they're talking to a human (which it can), where the TI-99 speech synthesizer voice absolutely can not.

Does that mean it's intelligent?

To me it just means they can brilliantly fake human conversation - the original design goal of Large Language Models.

It's really easy to tell if you're talking to an LLM if you ask a question that requires actually knowing things, not going for the first search result of a tool call or whatever most popular answer was embedded in the weights.

For this reason even the most sophisticated models still require system prompts, skills and all that other crap.

Most people are already at the anger phase, not at denial anymore. Get on with the times.

Why does this matter?

> Not if an LLM over chat can fool most people they're talking to a human (which it can)

I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing. I really can't understand how this gap persists; but then, there seem to have been at least some people who couldn't sniff out ELIZA, back in the day, too.

>I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing.

That's mostly true for longer LLM output with all the sycophancy / LinkedIn bias thrown in.

Make it casual conversation or comments, and give it instructions on appearing casual, or even better kill the censoring and fixed-prompt (with an open model), and it's orders of magnitude more difficult, unless if you suspect it and try specifically tailored prompts to sniff it.

There's no shortage of people obliviously discussing with AI bots in comment sections.

At the bottom of this very submission are a bunch of dead comments that are very obviously LLM-generated.

And that is proof that all LLM comments are easily found out? Also, easily found out by average humans? (this is not a average forum here)

Ok, you and I can easily spot LLM text. So what? The Turing test has still been passed, as is clear by people falling in love with ChatGPT, not believing something is AI, and by continuously claiming this or that is a bot.

People, many of them at least, cannot make this distinction anymore. You can, I can, but people as a whole are having problems with that.

>You can, I can

Even this (assuming it's even true) will likely not be true in some near-term future.

>continuously claiming this or that is a bot

I see it as a contemporary form of religious thinking. Like (say) pilgrims seeing blood on a statue of the virgin, plenty of people are now seeing the hand of AI in everything they read. If you want to see something hard enough, it tends to become magically visible.

The most interesting part of your reply is that you're not challenging the claim that the Turing test has been passed. I think it's a given, by now.

Sure. Of course it's been passed.

By definition, you won't be able to tell the ones that are fooling you apart.

Many, many people became friends/got romantically entangled with GPT-4o, to the point where OpenAI struggled to replace it due to user backlash.

Most users aren't very critical of the output. They just want a sycophantic ear, and 4o was perfect for that task. It's not _good_ but there is high demand for it.

https://arxiv.org/abs/2503.23674

From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"

Can it? I feel like I instantly recognize if I am chatting with an LLM or a human

"Feel" is doing a lot of work here.

You can recognize 70% of those (true positive rate) and still have a false negative rate of 30%, while thinking you got 100% of the AI ones!

The problem is that you'd be oblivious to those you don't recognize.

I was quite surprised on how difficult it is to tell when chatting with an uncensored LLM a friend is running (it's too big to run on any of my computers but he got some B200s). You can input your own "system prompt" to make it behave like a normal internet user and the prose writes very similarly to internet comments with none of the LLMisms from ChatGPT, Claude, Grok, etc.

Emphasis on "feel"

Can't even tell if you're real or a bot by reading one comment.

Great times!

I also believed that, but seeing qwen 27b overengineering solutions in a bit too familiar way in the article, I started doubting that.