Have you looked at using oMLX?

https://omlx.ai/

Would second this, I switched to oMLX I get ~75 tok/s on Qwen3.6-35B-A3B-4bit on a 48GB M5 Pro