DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
when are we going to stop pretending these benchmarks have any meaning?
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
Yep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.
global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
the work does nothing or causes net harm.
You’d expect that, wouldn’t you? But, alas…
https://en.wikipedia.org/wiki/Commodity_fetishism
"In Marxist philosophy..."
Well, that's about the same validity as "In Western astrology..." or "in flat earth theory..."
Would you care to discuss the topic, or just throw grenades? Surely you can come up with something more substantive than this
Ok. This requires the notion of "intrinsic value", which I believe does not exist (all value is subjective), yet is a foundation of all Marxist theory.
Buddy, admitting your thought processes forcibly terminate on pre-programmed keywords isn't a flex.
Why "terminate", Marxist philosophy is a legitimate topic, deserving to be studied. Like a rich sci-fi lore or a history of Tarot magic. Deep, fascinating, and wrong.
[flagged]
Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.
People already started using contributor API, and your input is irrelevant.
I'm using AI to build things I wouldn't (and/or couldn't) have built before.
That's the opposite of parasitic.
[flagged]
Don’t you have some looms to break?
And the unabomber has entered the chat.
I’m retired so it won’t be replacing my labor :)
The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?
[flagged]
Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
But is the score really reflective of the quality or are both models benchmaxxing?
Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.
Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
how much of it is from reallocation of staff to ai training and labeling