I think it was on qwen 3.5 9B or thereabouts that I saw a ~30% decode rate improvement on my M2 Mac vs MLX. That’s probably the best I’ve gotten.

Keep in mind though that this is also sometimes use case dependent. Off the shelf implementations are generally pretty good overall, but can have pathological behavior on specific workload shapes you care about. So do this as a somewhat later optimization, and particularly when you see performance characteristics that don’t seem to make any sense.