Dwarkesh recently pointed out [0] that these guys are almost forced to spend most of their compute on training instead of inference. This is because they need to maintain the appearance (which may also be the truth) that future models will make current models obsolete and be much more valuable.
Completely fixed-function HW can't be used for training, it's inherently a statement that "this model is Good Enough and we are now gonna start just extracting its value instead of extending it". So yeah it's an inference moat but it's not a growth moat.
Makes perfect sense for a company trying to get into the compute business, not companies who wanna be in the creating-ASI business.
Still, I guess/hope they have teams doing it in-house anyway. Just not something they'd wanna make a huge amount of noise about, it doesn't look good for To The Moon valuations.
[0] https://www.dwarkesh.com/p/why-compute-might-get-10x-more-ex...