the measurement breaks because the detector is probably also AI, and at that point you are just watching two language models argue about who wrote what