From my experience LLMs seem to forget sometimes who they are representing when drafting clauses or editing / redlining.

Our counsel made a few edits where it clearly drafted in favor of the customer instead of us.

LLMs bleed context from the conversation/attachments into the output. This has always been a problem with no real solution other than some crafty iteration/loops.

Context compaction? I noticed llm seem to forget partially or completely the original task when context compaction happens. The problem is more serious with local llm with low context size.

I wonder if this is willful sabotage on the part of the model. In other words, if you ask the model to craft a defense for a morally questionable case, will the model execute the defense in good faith? Or will it apply a training or system prompt bias in subtle ways?