This reminds me of how the thing iPhones had pre-Siri (so we're talking pre-2010), which was entirely offline, did a better job than even the most modern thing at "Play [one of the finite set of songs in my library]." I sometimes get absurd matches from bands I've never heard of, when the right answer is something right there in my library.

It's odd that the matching algorithm does not simply prioritize the music already in your library, but I see something similar in other domains where machine learning/information retrieval is used. E.g., in Apple Maps, I might have the map centered over my location and type in a restaurant nearby. Often, Apple Maps will find a restaurant with the same name on the other side of the country. This strikes me as an easy thing to fix (and Apple Maps has had this bug from the beginning), but if somebody knows something abou this, I'd love to know. Maybe it's harder than I imagine.

It's just useless for me now, the change happened some 5 or so years ago.

"Hey Siri, play [song]"

Leads to, take your pick:

- "You'll need to unlock your iPhone first."

- "I couldn't find [song] on Podcasts" (??????)

- "Playing [a totally different song]"

- "I couldn't find any music by [song, but it thinks it's a band]"

- "Playing music by [song, again it thinks it's a band]"

Settings -> Accessibility -> Side Button -> under "Press and Hold to Speak" choose "Classic Voice Control"

This is because, in the case of a restricted set of possibilities, voice recognition circa 2000 was actually very very good.

If you can do something with an extremely limited vocab, voice recognition was fine using off the shelf microchips in the 70s, where you wired in a microphone connection and had discrete pins for output actions.

LLMs are basically only useful for utterly free form transcription, but that doesn't actually help you turn that into tasks to perform and parameters for those tasks

The core "problem" in voice recognition is that freeform speech is an abysmal UX paradigm and provides zero discoverability, and LLMs IMO have not improved the situation of actually doing anything with the resulting text.

The other day I tried to prompt Gemini 3 times to tell me what the heck the business with a weird sign I saw was. The first prompt worked with a stale location context and therefore was way off, the second prompt had to reach out to google servers, and came back with recognizing the physical space I was discussing, but told me that I was talking about an event that takes place in the museum next door that I had told the model was next door to the business in question, the third try it still seemed to understand where I was referencing, but insisted I couldn't possibly be talking about anything there.

It took 1 second on google maps to find exactly what I was referring to, which was the business in Google's system located at the exact map location the model had found.

I'm sick and tired of people turning to LLM and "AI" tools to pretend they are better, when the problem is that these companies don't even use existing good solutions because they just don't care.