The interesting part is that the detector doesn't need to identify a specific token choice. It can look for a small statistical skew across many choices. That also explains why the approach is fundamentally probabilistic: paraphrasing, translation, or human editing can dilute the signal without necessarily removing every trace of it.