Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.
Actually I can answer my own question: we know that they have had previous access to the tests because they’ve run older models against the same benchmark.
I wouldn’t put it past a company like OpenAI with a long history of lying and being deceptive to record the tests and benchmaxx ARC. They have trillions of dollars of incentive to cheat any way they can.