If this were the case it would be necessary to send the entire model weights in response to every request which would be a bit inconvenient.

Hmm, could one instead of sending the model weights, send like, a merkle tree root for them, not exactly specifying the model, but at least demonstrating that the same model is used each time?