This has been pretty well studied. Pangram is about as good at an expert human with a 98% detection rate and <2% false positive, for flagging AI generated text.

https://arxiv.org/pdf/2501.15654

I'm open to being convinced, but the study you've linked me is limited specifically to professionally written, proofread journalism from 8 publications over a span of a little over a single year.

It's not exactly a wide breadth of varied text for such an all-encompassing task.