> GLM 5.3 Flash at Q4, AA score 57

That AA score is for the original model only

Then take DeepSeek V4 flash with AA score 52. Runs unquantized on 2x DGX spark with 1M context.

Or Qwen 3.8 27B, AA score 52 (which is utterly insane given the size of this model); I have been testing Qwen 3.8 27B since a week now, as an intensive GLM-5.2 and Opus 5 user - I can say that I just can't believe my eyes i.r.t. to how good this model is.