> This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words.

The specific issue of Google is that they are using an underpowered model, not fit to task, and much prone to hallucination than either OpenAI or Anthropic free tier offerings.

Google should at least match the frontier labs at the free tier (with some limit; after that, degrade quality), ffs

You're asking for something unreasonable. The number of Google searches per day is enormous and they haven't even been able to roll out AI overviews to everyone yet (they're missing in a new Firefox profile I just created). I wouldn't be surprised if the free tier frontier models cost over 100x more to serve than the AI overviews.

So then they should be pickier about when they show results or which model they use based on the question.

Nobody asked for an LLM response for every single search.

They used to detect certain types of queries and offer direct answers when the query matches. In my opinion that’s how Gemini in search results should work.

The specific issue is that search has become so bad that they think an LLM that gets answers wrong half of the time is a valid alternative, or, in fact, the "future" of search. Then they shoved that "alternative" to users with no way to disable it.

The less specific issue is that Google has no internal incentives to produce products that are useful to customers.