I’m not sure I follow their point. By going with structs of arrays you’re already getting your max SIMD bandwidth 80% there even with the most naive implementations. Learning which math operations are hardware SIMDable gets you another 10%. It’s the last 10% where you rely on math expressions (eg. trig identities), restricted pointers, and ternary “masking magic” where the author may have a point.

A recent example: https://glouw.com/2026/08/14/Ensim5.html