The no-reasoning version scores 35% while the low reasoning one scores 17%? What?

I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".

It's simulating the Dunning-Kruger effect.