Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day.

The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).

But the memory bus speed is fully committed when generating tokens or thinking.

Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.

The epyc Venice SP7 socket is apparently 9324 pins

https://x.com/tomshardware/status/2066846693778510331

We are going to need a bigger boat.

https://www.cerebras.ai/