Why assume they are not already making up the compute costs for smaller models?

Enough for such a massive reduction? If yes that’s really impressive