I run mine on an M5 Max with just 48GB of (V)RAM, and it fits nearly twice in Q4. Works perfectly. I'm kinda glad I didn't spend the extra $2400 to get 128. We don't really need more... and that's a good thing (tm). God knows I thought about it in store. But I thought... maybe this year will be the year of the local model? Maybe soon we won't need that much RAM? I was right.

The fact that it runs at 15tk/s in power saving mode, and 30 in perf. mode blows my mind. I can run the model in the background, coding something for me in OpenCode, hosted in LMStudio, while doing something else. What a world we live in.

Having something close to human intelligence (at least for reasoning and code), running on a laptop, is amazing.

Have you looked at using oMLX?

https://omlx.ai/

Would second this, I switched to oMLX I get ~75 tok/s on Qwen3.6-35B-A3B-4bit on a 48GB M5 Pro

For day to day LLM experimentation (and even some business use cases), I'd say Apple Silicon would be first choice for me.