These results look pretty good, given the smaller model size and the GLM family's historic robustness. Cheaper than Kimi and more robust than DeepSeek. The question in my mind is if you're going cheap, are you going to stop here or go all the way down to DeepSeek Flash?