why wouldn't they benchmark the accuracy against jev too?

Interestingly, I've seen better performance and similar cost to whats on JevBench

Is this an astroturfing account?

They say they are only benchmarking public models in the blog.

Also, section 2.3: https://typesafe.ai/legal/mca

Hmm, so actually I thought it would say that it's not permitted to benchmark or compare to other products, but I can't find such claim?

It does say "develop (or to facilitate the development of) a similar or competing product or service", but I think it would be a long stretch to say that's the case if they would just publish benchmarks. Microsoft legal department might disagree.