I also no longer trust benchmarks on this one. When the context gets a bit longer and the problem harder low quant models often produce worse output for me. Sometimes they even loop.

Interestingly different formats also often behave differently. GGUF unsloth is so far the best for me.