This will be roughly on pair with Kimi K3, but using a third of its parameters.

Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third.

Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.

Congrats def in order but as usual the proof will be in the pudding of actually running the thing.

GLM 5.2 has token efficiency problems. It's not a stupid model, but it takes a lot of "thinking" to produce not-stupid results. ("But wait...").

Which makes its pricing deceptive.

I tried to get by through the month of June on just GLM 5.2 and it was ... fine-ish for about two weeks. But the provider situation wasn't ideal.

Kimi is a great model but it was clear from the start they achieved they brute forced that performance through scaling. The frontier models K3 compares to are rumoured to be smaller also. GLM on the other hand is way ahead in perf/parm but severly compute bound. Now once GLM can scale up or Kimi optimizes the training more, that gonna be fun times.