Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.
Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.
100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes
Is anybody tracking these quality changes? All I've seen so far are accusations (quite a few at this point) but not really any actual data.
I don’t understand how there isn’t a website out there tracking this stuff already.
https://modelregression.com/
Haven't looked into how accurate the page is, but the list of regressions on the bottom looks terrifying, at first glance?
Yes, it looks like regressions are frequent, but sometimes performance goes back to baseline quite fast.
how the hell do we even track that? and before someone says...
—"Benchmarks!"
...I'll tell that they can be gamed so easily, and they are on a consistent basis.
Sure they are, but do you think they are continuing training to improve a model after release without bumping the version number, presumably only to game the benchmarks?
In a Codex subreddit there is a bunch of stats.
I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence!
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").