I have a similar calibration workflow, too [0]; but I am under no illusion that my personal benchmarks remain secret.

Once we send code & prompts to the providers, it can no longer be considered private.

[0] And came to the same conclusion as you did: Claudes were better but not by much: https://news.ycombinator.com/item?id=48654635