It's not the word I'm taking issue with, it's the concept. From your post you are surprised that the model makes mistakes and does not recognize things, and these are only surprising if you imagine the models to be thinking about things.

We can argue about words like thinking and intelligence all day long, and so can an LLM. But at the end of the day the LLM is only mimicking the processes you or I use.

Last week I tried out gpt sol 5.6 and asked it to count the letter r in a massive string of letters, without using an external app. It succeeded until I made the garbage sentence suitably large and then it consistently, confidently failed. Each time the "thinking" showed that it was teething to find a "gotcha" each time. "Ah, the first time I forgot to count the letters in the instruction itself" etc. at no point did it just understand that it had miscounted. It seems to be incapable of considering that it just made a regular mistake. Even a six year old child would just try again the same way and end up with the right answer