The benchmark comparison to other models was suspiciously missing. The self-comparison is a good representation of progress, but it is light years behind frontier models. Speed is great, but wrong is much worse than slow, IMO.