It's not "not quite true", it's literally true because alternative architectures like RNNs and Mamba fully update their own internal states, whereas transformers only append to the context.

RNNs and Mamaba do not update their weights, but you could hypothetically scale the internal state to be as big as Fable's and GPT 6's parameters.

At least one mechanism to update transformers' internal states already exists, there is nothing stopping anyone from performing backprop after every session.

It just has big technical and economic challenges. But I expect advances there. There have actually already been big advances, though done in bulk fashion (RLHF).