Now that is a good and interesting question! Hopefully a "no-moatist" will share their reasoning.

Because they're distilling frontier models and that's a lot faster and cheaper than training a frontier model from scratch?

I am not a "no-moatist" per se but one can argue their might be a plateau to how good a inference llm can become. If this is the case the playing field shifts to context, tools and harness, which are much cheaper to build an compete on.

Ultimately, that's what I need to be convinced. No one has put forth a good argument yet.

Clever architecture --> Ok but OpenAI/Anthropic can use these as well and they also have very smart people with their secret clever architectures

Distilling --> Ok but distilling means you will never be smarter than the original. Furthermore, reasoning is now hidden by private labs and they have poison pill answers for distilling if they can detect it. They will be able to detect distilling better and better.

Cheaper electricity --> Ok this is cancelled out by their chips being much less efficient due to not having ASML EUV machine access.

So I don't see why fundamentally their training costs are cheaper over the long term.

I'm looking for a no-moatist to convince me.

Labor. Smart labor would be much cheaper I'd reckon in China than in the US.

How much advantage in costs? What % of labor is training cost?

Mercor, Tacit Labs, Handshake AI... I suspect companies like these play a big part in model improvements, generating high quality benchmark/task-focused data for training.

However, these do require educated, white collar, workers.