I'm with the author and others in this comment thread, speculating that effectively the balance has tipped to where humans are no longer the target audience of post training - other agents are. Whether it's through the reasoning / CoT, or whether it's in handing off to subagents etc, the focus has moved to agents communicating in "agent-speak" to themselves or other agents. And human niceties are just kind of, noise in the way of getting work done.
This will probably bring us to a cross roads where the folks that want to remain in oversight and control of the AI work will bifurcate from those that want to skate straight to the future where nobody looks at anything and outcomes are evaluated purely empirically.
Every round of models (plus all the secret tweaks) require new strategies to stay afloat as a human. My new tactic for Fable and Opus is to give them a line limit, both during planning and code creation. It os amazing how well that works for keeping them on task and avoiding premature optimization, pointless tests or any of those "robustness" ideas that are not planned or asked for.
Before I even look at a PR of sol I ask it to justify the loc. Quite often it comes back suggesting things it could simplify
Could you provide more detail please? Sounds interesting.
Is this a hardcoded limit or something relative to the input prompt etc.?
I agree, and I'm actually pro Opus 5 exactly for this reason. In my opinion, agents talking to agents is the future, and humans will move to a higher abstraction layer. So it's the right move to make for Anthropic.
Many of the complaints that people are having with Opus 5 are actually acknowledged and explained on opus's 5 prompting guides (https://platform.claude.com/docs/en/build-with-claude/prompt...)
It also seems naive that Anthropic, with some of the smartest people on the planet working on AI, would not know about this.
More likely they know about but cannot fix it without tanking performance. It also effectively pigeonholes them into coding at the precise time they’re trying desperately to expand into general office work. You simply cannot use opus to generate any text fit for humans
> outcomes are evaluated purely empirically
How does that look like?
There's a neutral judge waiting out there to objectively tell us what good and bad outcomes are, obviously.
Or, even better, let's have the AI judge, because humans are too stupid to think for themselves. Look at all the harm they've caused in the world. Let's have this purely neutral AI with absolutely no hidden human intervention, decide how to run society.
This is interesting take. How do you think the split will show up? Harnesses/models will be designed for human consumption, and those for machine consumption? I'm building getwhelk.com and it is a cool thought-experiment for me.
or these models move to background agents and we need new ones for the humans to talk to