I'm pretty sure the answer is that their chip is uniquely unsuited for LLMs, probably because the write speed might be horrendously slow. Most likely too slow for storing the context window in multi user workloads.
I'm pretty sure the answer is that their chip is uniquely unsuited for LLMs, probably because the write speed might be horrendously slow. Most likely too slow for storing the context window in multi user workloads.