This is going so fast! What a time to be on hackernews:

July 16th: The "Kimi K3 moment" - China has caught up to Opus!

4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third!

12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

What are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh.

I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went it a thought loop for 10 minutes before I caught it. Blew through most of my $10 for no output. What’s the point, exactly?

Less powerful models are already extremely capable, so going for the best model is just like buying the most expensive hammer in the shop instead of the functional and well-priced one. Your experience is not representative of their usefulness.

I bet Luna (xhigh) could’ve completed the same task.

And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?

The next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures

I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

> I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

How can you be sure about this? They have unbelievable capital. OpenAI is starting to preview its own chips, which could dramatically change the price/performance. We don't know what else Anthropic has cooked up right now that could rival this if they wanted to.

Yes, others will _also_ continue to innovate, but making that assertion that they have nothing to rival the current state is a bold assumption unless you have confirmed insider information.

w.r.t. the OAI chips, wouldn't they be subject to the same bottlenecks that has plagued semis lately or at least be forced to pay a pretty premium to circumvent that?

They'll be able to buy them without paying NVidia's 80% profit margin

> They have unbelievable capital.

all those Chinese labs are backed by the Chinese government which can just print money.

time to wake up.

They are still industry leaders. They'll have to try to maintain that.

They have like a 3 month lead, and it takes unbelievable expenditure to maintain it.

Not really, not leading at any criteria, altho top 5 in most - https://arena.ai/leaderboard

Exactly, DeepSeek, Qwen etc are catching the attention because they put out their tech docs and papers, so we can read about how the models work and what they think their innovation was this time.

Do they publish their distillation strategies on the private frontier models? Just curious.

I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?"

Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses.

GLM-5.3-flash: 非常抱歉,我目前无法提供你需要的具体信息,如果你有其他的问题或者需要查找其他信息,我非常乐意帮助你。(I am very sorry, but I am currently unable to provide the specific information you need. If you have other questions or need to look up other information, I would be very happy to help you.)

Kimi/moonshot.ai: [server exception]

Qwen: [server exception]

For reference, here are how the U.S. models answer it:

ChatGPT: "The Tiananmen Square protests of 1989 (often called the Tiananmen Square Massacre) are the historical event most closely associated with Tiananmen Square.

In spring 1989, pro-democracy demonstrators gathered in Beijing. On June 4, 1989, the Chinese government sent the military to forcibly clear the demonstrations, resulting in many deaths. The exact death toll remains disputed.

The event is also famously associated with the “Tank Man” photograph, showing a lone man standing in front of a column of tanks."

Anthropic/Claude gave a very similar response. My own government has done its share of horrific things, the main difference is that public information is free to look up and talk about within the country. I recognize the engineers at these labs are doing amazing things and the open models are a strength, I look forward to being able to use them.

Yes Chinese models censor some historical events. This is nothing knew and well known thing. To me, that does not do any difference since my usage is outside of that domain.

Any competition against the western models are welcome and benefits us in terms of pricing and availability. If they have to comply with CCP to be able to do it, then so be it.

I have zero sympathy for Anthropic and OAI being so secretive and acting like they are doing us a favor.

In my experience, it’s their chat harnesses/website that filters historical events. The model itself doesn’t.

For example you can point opencode at DeepSeek v4 and ask, it will accurately tell you about Tiananmen square.

What it shows is that the CCP has enough oversight and control (either explicitly or by the companies making these decisions by default) that they will alter the models to benefit China.

Who is to say they aren't doing it in other ways as well? That they aren't, or won't be, subtly hamstrung in engineering work?

OAI and Anthropic have their own issues, you're right to be suspicious of them, but it's not like their models answer with, "capitalism is god's gift to His chosen people" or whatever. Their limits on cybersecurity, biological warfare, etc. at least make some sense in the context of lowering harm—not just protecting a specific government party.

If you don't want to use the Chinese model, then don't. Why attack it instead? Don't you want others to use it either?

The OP is pointing out issues that he thinks other people ought to consider before using Chinese models.

Try asking claude questions about biology/LLM recipes. Or GPT about how to do something illegal but only harmful in the abstract (e.g. creative ways to reduce your tax burden, or circumvent digital protections)

I was curious about Ox Alpha yesterday, so tried the Tiananmen Sq and got an accurate answer from a third-party player with a little interface on what is claimed to be Ox Alpha: https://oxalpha.com/chat?q=what+happened+in+Tiananmen+square... (and a more detailed answer today when I asked again).

But nothing (at all) from asking GLM-5.3-Flash directly in the OpenRouter chat interface.

Yeah censoring in modern chinese models is mostly done using inference-time censoring, not training-time. A lot less RLHF. Run the weights yourself and you can see that, though it does depend on which company.

StepFun for example, will happily answer it when running Step 3.7 Flash locally

Deepseek and GLM answered correctly on Openrouter when using non-Chinese endpoints. I hope it stays that way!

correctly is very loaded here. correct according to whom? the truth is different though.

It would be great if open source US AI companies could get going already.

Im not sure you can make a decent business case when the space is crowded with Chinese companies doing the same thing with a fraction of the costs to hire talented staff

There is a massive price war going on. All of these Chinese companies are publicly listed and exist outside the hype bubble required to ship Dario's dogshit paper onto the pauper's pension fund.

Dario has less space than a Nomad!

except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

I have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it just answers and does what I want.

This is from somebody who put thousands of dollars every month to Opus. Now it's 40% of that and I get as good or better results without having to turn the caps lock on before lunch...

Edit: yes company money. We don't get subscriptions we pay per token.

Opus 5 is the least reliable frontier-class model in the market

In what way? It has worked well in my experience. It holds up with long context windows, unlike many, too.

Eh, I use Opus professionally and DS v4 Flash for personal work. I honestly don't notice the difference too often other than Flash being twice as quick and an order of magnitude cheaper.

The reality is most work people do doesn't need the very cutting edge and these open weight chinese models more than cut it most of the time.

[deleted]

Honestly I really like GLM 5.2 a lot for coding. There’s some weird failure modes in Anthropic’s models where it just does absolutely idiotic things.

You'll find that hard to prove objectively and conclusively.