Why would the expected performance be relative to other models?

are you really arguing that opus 5.5 beating the best of openai's models handily is worse than you expected? What exactly were you basing your expectations on?