any watermarking ai researchers who can explain this?
what if the llm should
- repeat something verbatim (important in a compaction prompt)
- there is just one correct order of tokens for a somewhat long chain (a certain sequence of control signals)
- provide a diff of 2 inputs
without punctuation or whitespace wiggle room?
how does the drifting work?
does it postpone the drifting and drift stronger later?
what if max_tokens is set to a low number?
in what way does this not affect output quality?