What I don't see in the comments: "I had a specific problem I couldn't solve with the previous version of this LLM. But the improvements in this version unlocked the solution for me."

What I do see in the comments: subjective improvement in text generation, possibly lower cost, some optimism about code generation, but some skepticism too.

I use coding agents. To me they are very useful. But what I spend on them isn't going to support trillions of dollars in investment.

I had two sessions this morning that prior fable and sol sessions were stuck on, where iterations just resulted in _different_ bugs. (One kind of tricky fe layout problem, the other was a backend refactoring that was complicated by trying to aggregate a couple prior sessions that crashed).

I summarized each into new fable 5.1 sessions, and both seem to have arrived at reasonable solutions that only need a few nits revised before they are commit worthy.

So the issue was in large context size of these old sessions.

Commenters here likely haven't used it long enough to give non-superficial reactions. The customer quotes on the release page are all about it solving new problems, fwiw. We'll likely find out in the next few days how it really performs.

But yes, we might end up hitting the issue of "most jobs aren't solving hard problems" increasingly. The bigger potential benefit is higher trustworthiness, reliability/thoroughness, and squishy human things; people will likely continue to pay large premiums for those. "Solve it well and save time, long term". Those can be harder to see on a benchmark.

We rarely upgrade our phones or MacBooks because the newer version can do something the previous one literally couldn’t. Often it’s the efficiency, speed, battery life, etc, combined, that lets us push the hardware further.

I get your point, but we can only have groundbreaking leaps once in a blue moon. That doesn’t mean incremental improvements aren’t useful.

What you were describing our products at the top or near the top of their S curve. That only works if a product has achieved a mature market that's big enough to sustain further product development. Apple might take a percentage point of market share from Windows, and Linux might take a 10th of a point, but nobody is suddenly going to find, or lose, a big chunk of the market.

The problem frontier LLMs face is that they are hundreds of billions to trillions of dollars short of finding that market that's big enough to sustain capex commitments and further product development. If they don't find something groundbreaking, they are going to have a very painful year next year, maybe even starting this year for some of them and their data center partners.

Anthropic and OpenAI can't afford to live in a world where LLMs are at or near the top of their S curve.

I think the trillions are built on expectations that your employer won't need to pay you a salary anymore.

I agree. But it seems like this site has become so radicalized that this measured take is now anathema.