Please report success/failure after each test. Asking me to read and compare 30 writing samples to get any feedback at all means I won't finish. Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses.
Please report success/failure after each test. Asking me to read and compare 30 writing samples to get any feedback at all means I won't finish. Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses.
I managed 0/10 - far worse than random chance
Not sure what that says about me or the LLM but I guess I shouldn't worry too much about watermarking ruining the outputs…
I did five, then gave up and just pressed A until I reached the end. I got 3/5 right.
Exactly the same here. 4/5.
Yea, I did two and then harrumphed in annoyance that I was expected to do all 10.
We must do something about this immediately! Immediately! Immediately!
https://www.youtube.com/watch?v=jLO7VrRij_M
> Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses.
Wouldn't this make it a worse measurement?
Well, the alternative is a lot of us spammed A to get to the end, so the data quality is already horrendous.
Yes. Did one - saw that I wouldn’t get feedback until I have completed all 10 (if at all) and noped out.
I believe this is to test the theory that people can detect watermarked text.
Teaching you how to identify watermarked text while the experiment is running would ruin the data.
Maybe detecting watermarked text is a skill to be attained. Not allowing proper feedback and training will not allow people to notice the difference on time, thus ruining the data?
Best practice is to allow a number (scaled based on complexity of task) of training rounds (with short feedback loops) prior to letting people loose on the regular samples.
That would be a different experiment.
That's fine. But the author's goal here is to determine if anyone can tell. So he's probably not going to ruin his experiment.
Detecting watermarked text is hard. Detecting AI garbage text is easy but then classfiying the garbage further is beyond us.
> I believe this is to test the theory that people can detect watermarked text.
> Teaching you how to identify watermarked text while the experiment is running would ruin the data.
But that's complete nonsense. If it's possible to teach someone how to identify watermarked text, then you've already proven that people can detect watermarked text.