Yeah, that's definitely the frustrating zeitgeist for me these days, which extends to LLMs.
I'd also emphasize that it's not just error rate, but the shape/distribution of errors, and our (in)ability to build control systems around them.
To illustrate, imagine if someone unveiled a car which was unambiguously safer in every statistical measure... buuuut some of its unsafety came from jumping the curb to chase and kill pedestrians, under circumstances we can't predict for reasons we can't diagnose.