So far I don’t regret buying an M1 Max device with 32Gb of RAM. The models available for it keep getting better (running just about okay for interactive use) and 400 GB/s of bandwidth is still considered a lot.

The models are currently improving much faster than the hardware and this doesn’t seem to have plateaued yet.

Cool! I'm thinking about a local set up. What's your usual tokens/second rate?

Not OP, but I’m running local models on a M1 Max as well with 64GB RAM.

It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B.

I’ve also used Qwen 3.8 27B but I get 10t/s on it.

It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.

That's so cool. I wonder if the regular M5 can run those models too.