Cerebras doesn't etch the model onto silicon though. They're basically just wafer scale GPUs. They're more flexible than etched silicon though because they can just run the next version of the model almost straight away.

Interesting. I did not know that. And still they got 15x speedup (https://www.cerebras.ai/blog/openai-gpt-oss-120b-runs-fastes...). I wonder how much additional speedup would be possible by really etching the model.