You have a misunderstanding. Watermarking does not bias the responses in any way. How is this possible?
Before: "He leaped at the chance" - 33%. "Jumped at the opportunity" - 66%.
After: "He leaped at the chance" - 33%. "Jumped at the opportunity" - 66%.
But if you refresh your response from Anthropic 100 times:
Before: "Jumped at the opportunity" He leaped at the chance" "Jumped at the opportunity"
After: "He leaped at the chance" "He leaped at the chance" "He leaped at the chance"
The second one is detectable as being watermarked.
davmre has a good explanation that's more in-depth.