While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
Do you mean Fable 5.1? Or Opus 5.5? I'm not sure what you're working on but for me DS 4.1 flash isn't nearly at their level. For the price it's obvious very impressive, though Luna 6.0 is excellent too.
The problem with benchmarks and proprietary models is that one day a model is best at doing X, another day that's not so sure. And anyway, we are not throwing the same X.
I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at hand.
Fabble 5.1.
care to share what exactly are you working on?