Could someone knowledgeable please explain why an AI-focused CPU is superior to a GPU-based solution with an ARM CPU driving the high-level operations, please, especially when the RAM isn't on-chip?

I know that for GPU's the model weights have to be transferred over the bus initially, but that only has to occur once for inference use cases, so is the Fujitsu system more about training scenarios? Or is the focus more about efficiency, as these are ARM-based cores with AI additions?

Honestly I think this is an HPC CPU that has been AI-washed. Then you might ask why not use GPUs for HPC and I think the simple answer is that Fujitsu just doesn't want to take on the effort of building a GPU.