The difference is memory bandwidth. The M5 Ultra that's coming out on 22nd September can do 1,200GB/s. The M5 Max you can buy today only has 614GB/s.
The difference is memory bandwidth. The M5 Ultra that's coming out on 22nd September can do 1,200GB/s. The M5 Max you can buy today only has 614GB/s.
So, that gets us to about where nVidia was with Ampere in 2020. Let's hope the M7 catches us up with at least Hopper.
While true, the news here is the size of the unified RAM. Nvidia only exceeded 256GB RAM in the 2025 B300 - 288GB. The B300 alone (without the baseboard/PSU/chassis/wiring/CPUs/system RAM/etc) is at least 700% more expensive. This enables large language models on consumer hardware. 1200GB/s is plenty for many tasks.
> 1200GB/s is plenty for many tasks.
This + due to the hardware being so prohibitively expensive, we're seeing software optimizations happening. Like that dflash2 stuff for example, or an LRU for MoE and all that kind of stuff.
It is and it isn't. Why are you comparing the m5max instead of the m4ultra?
The big deal to me is the number of compute cores for prefill tps, which is suppose to be 4x faster on the m5ultra.
It's my opinion that the m5 ultra is going to be a really big deal in terms of local AI accessibility. Flash sized models (~200-300b params) are going to be reasonably fast as long as you aren't throwing 40k context at it on each or the first request (ie, agentic harnesses).
Even agentic harnesses like Cline should move at a reasonable clip on m5 ultra. I suppose we will know sooner than later.
FYSA: Former m4 ultra 512GB owner and current 4x rtx6000 owner here. I upgraded because I needed more prompt processing speed and concurrency.
What do you use all that local tokens/second for?
The complaint is about prefill which is not memory bandwidth bound, it's compute bound. But they added neural accelerators for matmuls to the shader cores which should make prefill faster.