> even though I have faster models available I still go to Fable or GPT-5.6 90% of the time

What about all the things you don't currently use an LLM for?

If a specialized chip can run a model 100 times faster, you can suddenly use it for a lot of things at sub-second latency. You can write "make white transparent and add a red outline to x.png" instead of the corresponding imagemagick invocation and perceive little to no latency difference. You can hook it up to your browser and have it yank out all advertisements live, or tell it to highlight anything that might interest you, again, live. There's probably thousands of latent use cases nobody has thought of that would be enabled by a truly fast LLM, even a mediocre one.

I don't think an on-device model needs to change much; it's already quite general in its capabilities.

Exactly, I think a lot of people aren't thinking about it this way, they're imagining a faster version of ChatGPT. In reality if it was a frontier model running a these speeds it would change so much about how we interact with computers. It would be custom hyper specific software on demand.

Looking for a lamp in a specific style?

"Make a VR application set inside my apartment (based on all the photos of my apartment from my photos directory) with all lamps under $100 that fit Scandinavian interiors and could be delivered to my house before Friday. Place the lamp on the dining room table, allow us to: cycle through lamps, change time of day, and interact with all the lamps and furniture".

Five seconds later and bam you have this new piece of software that you'll use once and then dispose of.

> hook it up to your browser and have it yank out all advertisements live

AI: "I'm sorry, a security guardrail prevents me from performing this operation."

or, hear me out, advertisers can do real-time advertising based on hyper-now context