Ahh sorry, I see that you explicitly wrote that deteriorated performance was due to throttling.
Naturally, LLMs do not get fatigue. As such they also don't need to detect it. That is correct. But LLMs can see if their performance have fatigue-like deterioration and correct for it.
That case is even more trivial and has nothing fundamentally to do with LLMs. Capacity is added everyday.
I would not base my long term projects on what happens in the market based on the fact that a single provider needs to throttle to satisfy demand now.
Your assumption that the current pricing is not sustainable might be fair. Personally I believe the opposite, and I havde not seen any indication that inference should not continue to decline in price.
In particular, LLMs appear to be hyper commoditisable. So if Anthropic or OpenAI is doing pricing shenanigans, people will likely move on.
> Naturally, LLMs do not get fatigue
they are not human so human sensations don't apply. but I heard when you run out of tokens it's also kind of like fatigue, just less predictable