It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
Well, I guess it's "fast paced".
What exactly is "paced" in this context?
It’s the famous “flattening the curve” from COVID. But for LLMs. This release is not flattening anything.
The word "pacing" (especially in the phrase "pace yourself") to mean go more slowly (at least initially) didn't originate with Covid. It's been around as long as I remember i.e. at least several decades.
I guess I understand why we'd want to flatten a COVID curve, but why do people want to flatten the LLM development curve? Don't we want the opposite? Isn't the goal AGI?
There is a difference between wanting AGI (which not everyone does), and wanting it as fast as possible no matter the side effects and potential for vast harm. Homo sapiens is 300k years old, maybe it’s ok to delay AGI by like… 1 year if it meaningfully improve our ability to align the model?
Well, there's a tension because, depending on who you ask, AGI is how you cure cancer and achieve utopia, but also how you kill all life on earth and turn the solar system into paperclips
I'm good with the odds on those 2 scenarios. I believe humans could kill all life on earth without AI anyway.
I think the goal is different from what "we" want anyway.
This is a preexisting model being optimized. Its absolutely not some unexpected release after that blog post. I won't defend that blog post, but saying THIS release is proof they don't mean they are slowing down is just incorrect, this is a prime example of what i consider horizontal improvements
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
What does model naming have to do with pacing or not? This is a ~20% relative quality improvement on the frontier (fable) at ~40% of the cost, just 21 days after the last release.
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
pretty annoying topic tbh. You're just weaponizing this dumb blog post so anything released is now a contradiction. By your same logic, if all inference was served at 50% less power cost and the savings are passed on somewhat to the user, its also a contradiction of the blog post.
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
Well yes, If you magically found a way to reduce a models energy usage by 50%, the thing that would happen immediately after is a doubling of a training scale and of the test time compute assigned for a given budget. I don’t see how that wouldn’t be considered pushing the frontier. Given that scaling is the one thing that has been bringing us closer and closer to AGI, a sudden 2x increase in scaling laws would definitely not be considered pacing. You seem to see a contradiction where I don’t.
Again, absurdly obnoxious, you could frame them giving an employee a shorter walk to their desk as "pushing the frontier".
1/6 the price, if they're right – that's a logarithmic cost-axis on the graph that shows Opus 5.5 medium matching Fable 5.1 max...
Fable 5.1 wasn't that much different than Fable 5 though.
[flagged]
That's a great point, five minute old account.
What do you think is happening?
IPO soon
[flagged]
Who said that, why do you take their word as the literal truth, and most importantly, what does this have to do with a focused discussion of Opus 5.5?
the prompt said that, so of course the anti-anthropic bot took it as literal truth.
Dario never said they would be. Just that more and more code will be produced by LLMs. All his predictions were in fact pretty much right in terms of months and percentages, give or take small margins.
Maybe not replaced exactly but they won't be manually typing out lines of code anymore. I haven't written a line of code in like 6 months. I review PRs, write prompts and tickets, check CI output, and get frustrated when the magical code machine stops working or I run over token budget
I mean they are literally getting sued since 3 days ago for trying to coordinate a slowdown, there is a very clear reason why they cannot effectively self-regulate.
Occam's razor: they couldn't make any more substantial improvements.
A razor is a philosophical tool to help decide between options, in the case of Occam it’s a way to decide for something in a situation where multiple options have more or less the same level of plausibility to en your current knowledge. It’s a heuristic to make a “cut”. What are you shaving off?
OP is using the expression correctly.
> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.
> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".
https://en.wikipedia.org/wiki/Occam%27s_razor
The options are some complicated narrative that leads the company to hold back certain capabilities due to some mercurial strategy or this is a typical release and there is no hidden strategy to decipher.
The "pacing the frontier" claims that they are slowing down intentionally?
Oh, you’re saying they were responding to the GP? Not their parent comment? Ok, yeah, makes more sense
I am responding negatively to parent, which claims they are self-pacing, when the simplest explanation is that they just have nothing substantial to show.
> Occam's razor: they couldn't make any more substantial improvements.
doesn't sound like a razor at all
Everyone knows both labs have internal models which outperform the frontier. All releases are to match market parity and demand for spend, the rest of the compute is used for training. It's not worth arguing about this.
Maybe their internal edge dried up in the last months
then technically the internal models are the frontier
Are those powerful models in the room with us right now?