> This seems to be one of the core alignment problems to me. See also: Gastown, the agent management project that could only end up working on Gastown, unceremoniously and quietly set aside.

I did not know this! Any link to an announcement or autopsy of sorts (even if not by the initiator of that project)?

I mean, it was pretty expensive, wasn't it? A few tens of thousands of dollars, IIRC?

This is the closest I could find to a post mortem from the creator:

https://yegge.ai/essays/the-shape-of-things-to-come/

But the GasTown part is barely a single paragraph that I could not make sense of. Like, what’s the Opus “tic”? Why was it so fatal to GasTown? As someone who only ever accessed Anthropic models through other harnesses like Copilot, I have no idea.

I do think what he’s saying roughly resembles what I’m forecasting will be a likely future of software engineering: that it will evolve into crafting comprehensive, bespoke automated validation mechanisms which let you establish high confidence in the agents’ work without really having to look at it.

> Like, what’s the Opus “tic”?

The article's description gives some idea, but I presume you saw that: "the 'just two more things' tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself."

Sounds like he ran into automated yak shaving. But it doesn't really explain why he couldn't fix it at the harness level.

Speculating, if you're trying to build something automatic, then you want constrained responses from each task you assign the model, otherwise you can get an endless explosion of work. The "Change 'Add to Cart' to blue" challenge parodies this: https://opusfived.dev/

Tangentially, reading the rest of that post gives me the impression that the author might benefit from an intervention. "AI psychosis" seems like it could be a relevant label here.