I'm going to cherry pick one example where newer models are noticeably improving at least in my experience.
What is a noticeable improvement with something that struggles to read a message longer than 200 characters without missing information in the middle, may be a 0.000000001% improvement with a model that... almost never misses info in the first place.
How does it work then?
I'm going to cherry pick one example where newer models are noticeably improving at least in my experience.
What is a noticeable improvement with something that struggles to read a message longer than 200 characters without missing information in the middle, may be a 0.000000001% improvement with a model that... almost never misses info in the first place.