I could definitely image Apple embedding a kind of LLM-optimized FPGA: slow to load (update) an LLM, but blazing fast at computing tokens.
Who needs memory when your model is set in silicon ?
I could definitely image Apple embedding a kind of LLM-optimized FPGA: slow to load (update) an LLM, but blazing fast at computing tokens.
Who needs memory when your model is set in silicon ?
You don't an FPGA if you're taping out your own chips. But that is just a MMA accelerator with decent memory bandwidth. No secret sauce here.
I am talking about reconfigurable gates to implement an LLM in silicon, i.e. an FPGA...