Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.

Lead doesn't really matter anymore. I just ported a very old cuda library to rocm, so it can be run on MI300s. 2 years ago this would have been a nightmare. Today it was an afternoon.

Exactly. Coding for inference is solved. CUDA is no longer a moat.

It was being served for free. They were almost certainly being overloaded.

Presumably the efficiency numbers they're quoting are for the high concurrency state they were serving.

RAM was probably the bottleneck for the amount of context they were offering.

I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful

Ox Alpha was also serving 10T+ tokens a day for free.

When it first launched on OpenRouter I was getting nearly 70 Tokens/second.

Has there been any confirmation about what that model even is?

Edit: Ah:

> This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.

It's also in this very announcement, in the first paragraph:

> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.

> and it was running very slowly

... I'm at a loss for words here. It was being served for free. To the entire world.

GPT-5.6 Luna is also served for free to the entire world with a tokens per second rate nearly 10X higher.

> ... I'm at a loss for words here

No need to be so dramatic. I think it's great that they're developing chips, but the whole "RIP nVidia" claim was overly dramatic.

Are you really comparing chatbot to agentic/code work?

Why is Luna not free on OpenRouter? :)

Do you know how much traffic luna was getting vs Ox Alpha?