I could definitely image Apple embedding a kind of LLM-optimized FPGA: slow to load (update) an LLM, but blazing fast at computing tokens.

Who needs memory when your model is set in silicon ?

You don't an FPGA if you're taping out your own chips. But that is just a MMA accelerator with decent memory bandwidth. No secret sauce here.

I am talking about reconfigurable gates to implement an LLM in silicon, i.e. an FPGA...