My experience has been that anything less than a 4-bit quant has a tendency to go off the rails. There’s a threshold of coherency that is being crossed somewhere internal to the model I guess.
Try the same prompt with a larger quant (even if it runs very slowly because the model no longer fits in VRAM) & see if Qwen does better - if so, there’s your answer.