I thought from what I read about the Taalas approach, the model architecture and overall size couldn't be changed, but model weight values could be updated after for further tuning.

Not as flexible as Cerebras though. And I'd love for someone who knows more to clue me in to the truth.

Nah, Taalas was putting the weights into silicon as a mask ROM. Their demo chip was hardwired to serve Llama 3.1 8B, and could never be updated. New models, even new versions without any architectural/size changes meant new tape outs.

But in exchange, you get insane speed and great energy efficiency. I could see it being a great approach for basic "good enough" models.

They may have had a little flexibility by supporting finetuning via LoRAs.