> numpy can't really do like native GPU execution

I'd be interested to see where GPU code beats NUMPY's SIMD implementation, which is really