This constant change of behavior, the dumming down of models over time as they do different levels of quantization to save processing cycle etc. To be honest I long for being able to get locked in versions of models with known parameters so I'm looking forward to getting more and more open source models and long term being able to afford running our own so we have a known stable llm model checkpoint and not what feels like random.
Not to mention you could at that point burn/etch the weights into silicon directly, and have models as 'ROM cartridges' that could perform at thousands of tokens/second, enabling entirely new use-cases.