The same thing holds for speed. I could build a system that speeds up Fable on GPQA Diamond ~50%, while improving score, by literally randomly selecting between Fable and Gemini 3.7 Flash. (Solve time for Flash is 0.1min and 0.8min for Fable, with Flash having a better score.)
Hell, I could publish better score at 87.5% time reduction by having the router always pick Flash!
Like I mentioned, we omitted routing from the benchmark and evaluated with our strongest model modality. Also, randomly selecting between Fable and Gemini 3.7 Flash wouldn't preserve quality.
Mentioned where?
What do you mean it wouldn't preserve quality? I demonstrate in the comments here how it would on a saturated benchmark.
Mentioned in my comment above. Sorry, I thought you were talking about Gemini 3.6 Flash, not their new one. Yes, that's the reason why we didn't include the router in the benchmark to have a fair comparison of the harness itself.