Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.

I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.

Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.

Canceling my Anthropic Max sub when this ships.

yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).

Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.

Could you share more about 5x/20x? I missed that

20x related to the 5h limit only. Weekly seems to be around 10x, although they deliberately don’t give a number.

OpenAI is 20x on both limits

> Weekly seems to be around 10x

Actually no. 5x and 20x have same weekly usage across all models. Just ask their chatbot.

https://x.com/beydogan_/status/2095293596198957418

it's clearly wrong, think it's realistically closer to 1.7x

Sol easily outperforms Fable on every task I've tried it on.

That's not my experience and I suspect it's not most people's experience. Out of curiosity, what's the hardest task you tried?

For me something the likes of: design a CDM for integrating these 5 logistical systems, with full docs and examples provided for each, as well as modeled transports specific to our business. Prompt was of course much longer.

Both failed spectacularly. But sol's output at least contained interesting findings and some useful parts, as well as not being 20000 words of unbearable language.

I can't speak for others but I have a feeling you're in the very small minority with this take.

You could say Sol is faster and cheaper and that's true. Outperforms Fable? Impossible to believe without hard evidence.

I dont think that feeling is entirely useful.

Because Claude doesn't allow third party harnesses on their subscriptions I doubt the majority of signals you're getting are actually that significant on pure model quality.

I suspect you're right on Sol not outperforming Fable; but i've not used Fable that much.

---

But, fwiw, in my custom harness between Sol & Opus 4.8 - then Sol wins by a ridiculous margin as Opus keeps claiming slightly wrong things with certainty much more.

This is not saying much. Opus 4.8 is ancient history.