We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
very few people comprehend - how much of an asteroid level event for western AI labs this is.
china has cheap abundant power, now they can make their own inference chips (which was supposed to be a chokepoint), their models yeah can be 6 months behind the frontier - but most people don't need frontier models - small models r more than enough.
my only wish was labs like Mistral would make their own inference chips or partner up eg with established / new chip makers or companies like Oxide.
For non residential consumers electricity is actually more expensive in China than in the US https://www.iea.org/reports/electricity-2026/prices
Strategic industries generally get subsidised/free power.
Perhaps more than the price advantage is the prioritisation in grid infrastructure. As a strategic industrial concern in China you almost certainly get easy access to transformers, grid connections, water etc which is a big bottleneck in the US.
While the North American models say 'No', the chinese models say 'Go Go Go'. I guess we'll see whether the anti-consumer wins over the pro-consumer.
From my understanding Chinese companies are pushing for open global cooperation on AI, as well as Meta.
Only the US sees this as a competition, new space race, Cold War, etc.
So effectively, trying to undercut their competitor to reduce its wealth and power?
After AI agents get good enough the real bottleneck will be power generation and political systems.
There is one player who might have a trump card up their sleeve: free power, in orbit. It isn’t over yet for the US.
Also, don’t underestimate data retention and such. Big Corp will never send their LLM traffic to China.
Thats a joker card. Anything in space is just extra complexity. And raw training data is no bottleneck any more too.
Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.
It wasn't a secret either. They blogged about it last month: https://z.ai/blog/glm-5.3-flash#:~:text=Serving%20at%20Scale...
[dead]