How could you establish what parts of the code was produced by an LLM vs updated by a human afterwards?

The LLM will output different results over time as the models get updated. Are we heading towards needing to retain a full prompt history that can be replayed against a specific LLM model version to prove what the output was for copyright purposes?