On the one hand, the Claude code thing was indeed a weird brag, because they took shit performance and made it suck slightly less. It's not in the same ballpark as getting a 2% boost in an already highly optimized system (e.g. Google search results page lets say).
But it was also a tale of simply needing to tell the models what to care about. They built metrics (I'm sure they vibe-coded them) and then told the system to optimize itself to those metrics. And it sort of worked. Of course most LLM code is not performant by default, you still have to prod it to optimize itself. You can't one shot it yet. But nothing about the trajectory of this stuff tells me it won't get there some day soon.