Given the huge amount of money being spent on AI chips in the US, what prevents US AI labs from doing the same level of software optimization? It could be a solve for some of the capacity constraints.

They have already been doing it for months https://openai.com/index/openai-broadcom-jalapeno-inference-... . OpenAI on their custom chip brought up lightspeed deepseek as experiment by using AI in the exact same way as this zAI blogpost. And the kernel optimization contests/etc have all been havily done through AI based optimization loops for half a year+.

> what prevents US AI labs from doing the same level of software optimization?

Because they don’t have to. Most of the time money would buy you newest and/or more hardwares so there’s low/minimal interest to optimize the code or approach.

They are making these optimizations. They publish these reports too if you care to look.

I would be extremely surprised if US labs weren't aggressively trying to optimise their stacks in exactly the same way. Any gain in performance or efficiency directly affects the bottom line as well as research speed.

They do, when Luna got 5x cheaper it was directly attributed to some unknown % inference optimization.

US labs are quite cut throat about dealing with stuff costing them money (inference). This sort of engineering excellence doesn't always feel that way because they are simultaneously quite lax about stuff costing other people money.