Maybe detecting watermarked text is a skill to be attained. Not allowing proper feedback and training will not allow people to notice the difference on time, thus ruining the data?
Best practice is to allow a number (scaled based on complexity of task) of training rounds (with short feedback loops) prior to letting people loose on the regular samples.
That would be a different experiment.
That's fine. But the author's goal here is to determine if anyone can tell. So he's probably not going to ruin his experiment.