GLM-5.3-Flash is actually cheaper than deepseek and better than deepseek but no one is talking about yet :)

It's actually slightly more expensive ($0.50 vs $0.48), but there's a temporary 50% discount.

I've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.

It's interesting that OpenCode Go is treating it as 2x more expensive than DeepSeek Flash, even factoring in the 50% discount

OpenCode Go is probably using quantized down DS4Flash. They outsourced to 3th party providers to keep the cost down, and being able to provide that $30 value (instead of the initial $60 > $15).

We saw the same issue with GLM 5.2 when they still published publicly who the providers are on their website. Most ran FP8 but one was doing FP4, so you had this issue where one moment you had the better FP8 and another session you had the FP4 provider.

You can check the internet archive, it was in the FAQ part before they hide/removed it. So if you looked up the providers, and the published quants, yea, ...

Given that a lot of complaints are coming from people that felt OpenCode Go Flash feel like a step down compared to old OpenCode Go/DeepSeek API directly, it smells of a quantized down provider is mixed in.

OpenCode Go is becoming less of a good deal by the month. I pretty much only use it for mimo 2.5 pro now, and everything else is either ollama or openrouter.

Go has API pricing + this weird scaling of how much is it worth. Some models get $60 of usage, some $30 and some $15 etc.

Not in my experience. Tasks that would normally cost $0.08 on DSV4-Flash have cost me $0.30+ on GLM-5.3-Flash. These costs are after Deepseek's recent increase. Also GLM-5.3-Flash is so slow compared to DSV4-Flash. I would be fine with GLM-5.3-Flash if it was cheaper and at the same speed as DSV4.

I use DSV4-Flash on Max through Deepseek's API. I have been using GLM-5.3-Flash on High through Openrouter which I thought had a 50% discount. I must be doing something wrong for the costs to be off this much.

I've been using it quite a bit too. My main complaint is that it can be really slow sometimes — like, really slow — and the speed feels pretty inconsistent.

z.ai is using all Chinese hardware for flash: https://thenewstack.io/glm-5-3-flash-chinese-chips/

There are other providers with much faster inference, like BaseTen at >100t/s: https://openrouter.ai/z-ai/glm-5.3-flash#performance

Does Chinese hardware mean fabbed in China or designed in China and fabbed by TSMC?

How do I find out where the openrouter model providers' servers are located?

If you click on the provider name, the panel that pops up shows a "Region" value. Not every provider lists their region, however.

I think the region is just the HQ of the provider. So z.ai's region is Singapore but it's quite likely that their servers are actually in China

I don't think that's right, or if it is, OpenRouter has incorrect data. Several Chinese companies (headquartered in China) have Singapore listed as their region on OR. And some companies, like Alibaba Cloud, have multiple regions listed.

I'm happy to be proven wrong, but this makes me think that the region is where the servers are, not where the HQ is.

I couldn't find any article that states z.ai has a data center in Singapore. There are stories of their new 1 GW data center in China though. Also, openrouter lists HQs on their providers page which matches the regions. https://openrouter.ai/providers

Interesting. I wonder where they're getting their data from then, because they list Z.ai under Singapore, but everything I'm finding says they're based in Beijing. Same with MiniMax.

Many Chinese AI companies have their HQs in Singapore because of US sanctions. If I remember it correctly Manus is also based there. But it's a Chinese company through and through as we know from the events that unfolded after Meta's buyout attempt.

It’s cheaper sure, but it’s very slow. It’s not a drop in replacement

I think we don't have a good draft model for better speculative decoding yet (e.g. DFlash 2). Once we do, it will be faster.

It very well could be faster, but right now it isn’t.

It is a slow for me through z.ai; it does not feel 'flash' at all. But then neither did the new DS Flash. I think they were getting hammered.