I have not seen that kind of significant drop with GLM 5.2 yet? so curious why you think it will happen for K3.

This is a very large model. Much larger (3x) than GLM. The resources to run it are very expensive.

There's been a price war going on openrouter between providers of GLM 5.2. NovitaAI, DeepInfra, and StreamLake keeps underbidding each other in waves. Yesterday evening both input and output $/M was ~$0.3. Output was especially cheap.

I had the opposite experience. I bought the $100 monthly sub from Neuralwatt last month because it was the only economical provider for GLM 5.2. They raised their rates halfway through, and it simultaneously became too slow to use.

I just looked at DeepInfra -- I've got an account there already etc -- and it's at FP4 quant. How much that effects the quality of inference for GLM, I can't say. I could see using it as a backup when other things run out but don't think I'd trust it.