Good benchmarks are costly to build even for mid-large corporations. And once the benchmark is used on models you really can’t tell if the problems would be scrapped for training

I was speaking in general, not just about AI.