There is something deeply wrong with Astra. I can’t quite put my finger on it. On the one hand it is a lot more knowledgeable, which makes sense since it’s a larger model. On the other hand that knowledge doesn’t reliably translate to intelligence or insight. Certainly tasks like 3D modeling it does extremely well. Other stuff like complex coding problems in an existing codebase it stumbles more often than not. This morning it ran around in circles. It implemented a feature, then convinced itself that it should have followed “proper TDD”, deleted all the code it had written and wrote 8,500 LoC of unit tests. At that point I was down to 35% of quota so I stopped it and gave the task to Opus 5.
Really weird model. No idea how it did so well on all the benchmarks.
Which benchmarks? Only the ones OpenAI cherry-picked.
It debuted as ~same score as Sol on Artificial Analysis. People couldn't accept it so they had to change the formula.
The model is a big step forward only in desktop use and 3D. That's impressive, but for software engineering, Fable is still in a league of its own.
That's exactly the kind of behaviour I've seen, unbelievable amount of unit tests, and revisiting and revising the same code over and over again. If I was cynical, I'd say almost like it was deliberately trying to burn quota, even after I told it quota was getting low and to move onto actual physical testing.
was it maybe over-quantised to reduce costs?