This feels to me like a lovely example of both high-quality systems design and information theory. From my perspective, one of the more abstractly interesting thoughts surfaced by this write-up is the degree to which you've "encoded" a highly complex set of relationships through use of metaphor.

That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.

Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.

Thank you for sharing and the care you put into writing this!

Thanks for the thoughtful response. We too were surprised by how deeply we could take the model of the teaching hospital to software engineering. Whenever we think to expand the system in one dimension or another, the teaching hospital model seems to have a nearby analogy.

I will confess however that some people internally find the model confusing. For example, one user couldn't remember that to get an issue actioned, they needed to put it into the "waiting room". They wanted to be able to action their issues without learning how complex medical systems operate. There seems to be more work on our plates to make this resonate with all users.

Now imagine what it is like to run an actual hospital with real people whose lives (or the lives of their loved ones) are often at stake, with underpaid and overworked staff and with messy biological creatures as the subjects instead of bits and bytes. If anything this whole exercise should also give you a much deeper appreciation of the people that feel themselves called to help others.

I'll be very interested to read any followups about the part you described at the end about getting it to do more "teaching," both from the perspective of how we can use this to help newer devs out but also from the perspective of helping ourselves understand the software being built.

One thing I've been worrying about recently is the notion of comprehension debt and making sure the humans can still understand the system (so as to be able to do incident response or something if the LLMs are down). The slower, more methodical workflow described here should definitely help by allowing checks like proper docs whenever a piece of the architecture changes for instance, but I'm curious what else could be done to help the humans understand as much of the implications of what the AI is doing as possible.

Edit: having the code author and reviewer be completely separate like you described below is also probably pretty helpful for making sure changes are comprehensible from the outside.

[dead]

I had a similarly useful analogy of the legal system emerge, in a similar way. I think there's a great analogy between policy in the legal system and policy in software engineering, and they have some awfully good (and very, very historical) ways to think about e.g. amending, repealing, and adjudicating things based on those policies.

i think role-play and utilizing the full power of language, stories, character, and narrative will unlock very sophisticated use-cases and in general a new dimension to agentic systems much in the way you're describing.

probably what will work best in the future is specific training or fine-tuning against curated datasets of narrative fiction and/or texts in general? not really sure.

I feel like this is something we do with humans as well. Things like "scrum", "sprint", "the clean coder", etc.

yeah this is really cool. i've been thinking about sort of a mirrored idea, where you end up with wizards, mages, clerics etc and fantasy terminology used. idk if it would be as effective as this is, but maybe it would be more creative somehow?

i kinda want to try to build this off of github. it's essentially just an event bus / message queue that workers (agents) tap into.

The metaphor from the title does not speak to me. Treat bugs like patients. How. By making bugs "feel better and be stronger"? I'm sure they mean "treating bugs like illnesses".

Exactly. It's the app that is being treated like a patient.

> it’s consumed just over $135,000 in Claude tokens

Ironically also just like a US hospital. Go in and come out with a six-figure bill.

In this analogy we're basically the insurance company. "No, we won't pay for that on the current plan". "That's a preexisting issue, we're not going to do anything to help with that"

I don't know if the US insurance/medical stories are a joke or true at this point. You guys live like that for real?

> You guys live like that for real?

"live" is doing a lot of work there.

I can think of nothing I'd like to use less than a database or filesystem vibecoded by Gas Town-flavored psychosis. Roleplaying with LLMs is not the secret to producing amazing code.

Seems to me it's just research and experimentation.

A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?

I posted in another threat about this: I am seeing a lot of people building their own little bespoke factories. They introduce endless quality gates until it slows development down, and then they add more agents to decide when to run certain actions, and on and on. The end result from what I've seen and personally participated in, is that it often winds up providing negative value in the software development lifecycle. It creates a whole lot of heat, but IMO is not helping the teams utilizing them to ship value any faster than they would with a more limited setup.

I just finished ripping out one of these “dark factory” setups. Removed about 70k lines of code and 750k words of generated documentation. For what is effectively a 5-screen app.

These software factories are just s/code/specs/g. 5000 lines of (agent generated and poorly reviewed) specs for 2500 lines of code. All because the CTO dreams of getting a few likes from Elon Musk.

Funny that we've automated meeting hell. Definitely a truism that we reimplement the org chart in software.

I've gone through this cycle recently. I think static lint/type checks are super useful, but agentic code review loops can easily go off the rails.

> not helping

It does! It helps getting promotion with tokenmaxxing.

lol, no doubt.

I think this is definitely a legit failure mode with this model. We’ve hit some parts of it. For example, a plan reviewer kept pushing a technically valid but negligible security finding through five review rounds and about $70 before a human closed it as won’t-fix. The "Research Department" ran every three hours and was filing around 8 new issues a day, more than we could review. Review nits used to spawn new issues of their own.

What helped us maintain a degree of sanity is that every stage records tokens, time, and cost in a ledger, so these show up as numbers rather than vibes. The fixes were mostly about giving a stage permission to stop. Plan review can now propose won’t-fix when the exposure is negligible. Research runs daily, with a cap on unreviewed proposals. Small fixes now ride along in the PR instead of becoming new issues.

The other thing is that the quality bar should match the domain. Our migration tools move customer data into a database, so our question was “would we be embarrassed to have merged this?” For small apps I’d agree most of this is overkill.

Fellow HNers: this is one of the authors of the post we are commenting on.

Please do not flag him into oblivion again. Sheesh.

This is Microsoft scale, I'd be surprised if it made any difference these labs. It's more likely it's plain managerial mishandling of the infra in chasing new profit heights.

I really like this writeup. I find the metaphor satisfying, though there's a reason metaphor receded into the bushes as a practice in XP. A great metaphor is great to have, but they don't always exist. Fortunately with organizational systems, things will tend to rhyme. For example, a role that "Verifies that all processes have been followed before anything merges" seems like it could fit many different systems. We're at this fun stage where people are just trying stuff and I love it. I hope we get more info from this experiment.

The next role: Insurance Rep.

Ensures all other agents are operating efficiently and within reasonable levels of token usage given the expected level of effort to 'resolve' the patient. In moderate to severe cases, may lead to patient defenestration or agent revolt.

Or perhaps, the author considered this role but found it typically costs more than it saves.

This is a good idea; so far we've only had a few informal chats around adding something like this. Right now we play the role of Insurance Rep ourselves by manually checking on our cost dashboards and investigating expensive treatments. (As did our Director and VP on the day when we experimented with using `fable-5` agents for everything.)

“patient defenestration” certainly conjures an image.

I am curious why so many of these systems are based on GitHub issues. Why not use a proper ticket system?

I use Redmine (lightly customized) and have found that to work very well. Especially for performing review / UAT of the work or providing response to the agents questions when it wants me to pick an option.

The underlying data model of a ticket / project management system is so rich and well suited to bodies of work.

Who defines what proper is?

Why are github issues not a proper ticket system? E.g. rust uses it, seems to work fine for them...

The simple answer is convenience. It was just easiest to build it this way when we started.

Rafi and I, who authored this post, will be hanging out here for any questions people may have.

Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?

I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.

No, we haven't performed any ablation studies yet - it's a good suggestion.

The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.

I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.

> One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.

I use a similar process for a local agent swarm approach (locally hosted Qwen-3.8-27b or Qwen-Flash-Next), used so far for personal-grade projects. Through tool and process accretion, the review stages are told to check both the work and reasoning of the implementation stage, via processing the pi.dev session log.

Through parsing the jsonl log, the reviewer sees the subagent prompt, the tool call sequence, and non-thinking narration along the way. That allows the reviewr to audit the implementer's process (e.g. were tests run?) and spot procedure violations or gross hallucinations. The reviewer also independently runs the test suite, so even a hallucinated pass is caught.

It's relatively expensive both in tokens and time spent running ideally duplicate tests, but the independent workflow has nonetheless caught errors that would very likely have been missed by a same session, same context review.

[flagged]

Yes agreed, using a system very similar without the hospital vibe :)

Could you fix a bug in this codebase yourself, without the help of agents?

Maybe trying this would be a good experiment? Say route 10 or 20% of any future issues to the “Chief” to troubleshoot and fix directly, and record those for further work?

This, with an extremist take on code quality via linting in ci, is no doubt the future, at least for maintenance and extending the interface kind of work.

1. How does this work with greenfield lifts where the scope and final vision are not yet figured out?

2. Sorry if I missed this in the post, but will you open source this system?

How much of this is a snake eating its own tail against vs the control group?

are any of these artifacts public and available for inspection/use?

Not currently, but open sourcing this has been discussed. Stay tuned.

The blog post links to an issue in the Sinai repo, but it’s private or just doesn’t exist?

Sorry about that. The repo is internal to Cockroach Labs for now. It was an issue about adding a progress rollup comment to decomposed parent issues. We’ll get the post updated.

Which inherent limitations did you recognize in the metaphor before commencing this research?

One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have "treatment" start on the cheapest possible model and see how it ended up, only escalating to more capable/expensive models as necessary. We rejected this idea because we suspected that it would end up in more time/tokens on the capable models to fix any problems.

In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.

[flagged]

I set different kinds of structures for different projects, Complete with setting and ambiance. It’s like agents work better if they are role-playing. It’s extremely disorienting and people with marginal mental stability are going to really have a bad time. What have we wrought?

Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.

I'd say that the first "mistake" we made was having it run in auto-merge mode. It was incredible to see what it could produce, and the speed with which it worked, but while the results seemed good, they were being produced at a rate that we couldn't human-verify. This is not to say that they were bad, but that we had no way to convince ourselves that they were good.

Once we started to review the code in more detail (and put the hospital into "human approval mode" to stop the momentum) we found one major thing that we corrected.

When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We've since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how/when the deferral should be actioned. This has prevented several items from falling through the cracks, but we're constantly refining what is a valid deferral and what should be fixed immediately.

There are several other things that we've corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.

> When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped.

Could you structure the DAG so that after each node that contains the work, you have one dependent node that verifies each distinct requirement was implemented as expected?

This way if a single requirement is dropped the system alerts you rather than it being silently dropped.

It makes sense intuitively that if a task has nothing that depends on it the LLM might accidentally attempt to drop it (even purposefully as an optimization).

I like this. Today the parent issue closes automatically once every child is checked off, and nothing re-reads the parent’s original ask against what actually shipped. There is a deferred-scope auditor that catches the deferrals an agent admits to. Your "verify node" would catch the ones it doesn’t, which is the more dangerous kind. I think it would fit naturally as one final child per decomposed parent.

Whoever is following the author of the post, this account, around, immediately flagging everything he says into oblivion

knock it off!!

they're getting auto flagged by hn's bot detection, since they appear to be using an llm to write.

Is a code comment and ledger the best way? Should the agent just fill out a form or something and attach it to the sub issue. This is how the hospital would work.

It often does that too, but we keep the ledger and the code comment as well, to ensure that if it doesn't get resolved by the sibling, that it's not lost. If the sibling does resolve it, it's removed from the ledger and the code.

[dead]

Maybe I missed it, but I didn't really see anything about the long term quality or maintainability of the code. All I see is agent agent agent.

It's true that we don't have long-term maintainability data just yet. We've just shipped the first product of this model to customers and likely won't have any detailed maintainability data for several months (and for good data, several years). We hope to author more blog posts on this experiment in the future.

Perhaps doing a random sample of the steps by hand will be a good way to notice maintainability issues?

[flagged]

If I wanted to build something like this, at least conceptually with the roles, where’d I start? My first guess would be to give Claude your blog post; but any other pointers to make it work reliably? Do you happen to have the system open source?

Great post. Thank you for sharing it on HN.

Modern best practices for a medical team (standard roles, responsibilities, consequences, processes, etc.) have evolved through trial and error, since the advent of civilization, to prevent costly human error.

In hindsight, it's not too surprising that the same best practices can be applied by teams of AI agents to minimize costly AI error.

It's still incredible. We sure live in interesting times!

You could also run them like an army, like an artist commune, like Valve (add virtual tables on wheels which they can pull wherever they want), like a lawyer's office, etc.

[deleted]

You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?

All that credit goes to Rafi. I believe that he either noticed that they were getting long winded in a code review, or suspected that they needed trimming after reading this blog post from Anthropic: https://claude.dev/blog/the-new-rules-of-context-engineering....

It was definitely human driven, but I believe that the agents did the actual trimming.

That is helpful. Thank you :)

To add a bit more: it was motivated by a general desire to not let costs spiral out of control, and when I read the Anthropic context-engineering post I wanted to apply some of the lessons.

The audit was agent-driven, but it worked from a rubric I came up with. For example: - the same rule restated in several places - narration about what other stages do later - history and rationale that don’t tell the agent what to do

One analyzer agent per skill file flagged passages in each category with a word estimate. Each also produced a “keep” list of things that must not change: commands, gates, templates, sentinels. Agents also did the trimming with some mechanical rules: no command, bash block, label, template, or numeric threshold could change, and we checked the diffs for that. The first pass cut about 18%, because it kept anything it wasn’t sure about. A second pass used an auditor plus an adversarial verifier for each file, and took another ~600 lines out of the six worst files.

The more durable result from all of this was a short style guide for editing skills. When an agent makes a mistake, the natural fix was “add a sentence,” and that’s how the files got that way. Also, long inline bash blocks moved into scripts with their own tests, and extracting them turned up a couple of bugs that had been in the skill file's prose.

One role I didn't see was patient advocate/representative? That might be another approach to non-convergence - "how is this going?" and escalation.

We actually have a /sinai-advocate skill, where a human can advocate on behalf of a stuck patient. We use it every once in a while when the labels get screwed up, or a workflow fails for some reason.

Great post!

One thing I've been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?

That's a very interesting idea. Sounds more like a "routine follow-up" in the medical model.

[flagged]

[dead]

I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.

Fable feels like overkill for this also.

We don’t run the most expensive model everywhere anymore. We started on Opus for every stage, then reshuffled several times. We landed on Fable doing planning and code review, because code review is the last gate and its bug findings were worth the most. Opus does coding and plan review, and Sonnet runs the nurse roles. The planner can also route mechanical changes to Sonnet. The plan reviewer re-checks that choice, and any rework goes back to Opus.

That’s tuning from the cost data, so it's not just totally based on feeling, but it's not a controlled A/B test. I agree that would be worth doing, and I think our experiment here is helping us make the case that it's worth spending more time and effort on running more controlled evaluations.

Create a “token guy” that others can consult with to figure out the best model to dispatch something to. Token guys role can be defined as an expert in the current LLM landscape. Not sure whose role this is analogue to in medical (somebody here said insurance agent… that is probably getting close).

Token guy’s role gets even better when you let them recommend entirely different hospitals for a given step in patient care (OpenAI models, random shit on openrouter, etc).

I have a “token guy” in my workflow and the dude has looked at underlying tool call patterns, cache use, etc and made all kinds of recommendations to the clients calling it all on their own and completely unprompted—just following a simple role.

I’d experiment with expanding their role into prompt engineer or some kind of specialist on that domain.

If the correct analogy for a team of SW agents isn’t a team of SW engineers why is that?

For example, we don't have a clear analogue for "triage nurse". What would be it? Oncall engineer for production issues; "product owner" maybe for the regular tickets?

Thing is, software engineering is still (relatively) young. And more of a craft than true engineering. So the "good practices" are _somewhat established_, but at the same time, not quite. Like, we have glimpses into what works and what doesn't, but not a universally-established workflow to follow. And AI is upending what we already knew; it's really not surprising that people are looking elsewhere for workable metaphors.

> For example, we don't have a clear analogue for "triage nurse". What would be it?

Triage nurse seems well defined in software engineering, perhaps just without a single title. There are certainly engineers that focus on triaging GitHub issues and then fixing immediate showstopper bugs and/or prioritizing the issue for further attention from other engineers.

That role can be communicated to agents by writing down that description. The risk with saying that the role is “triage nurse” is that you can’t be sure what other roles and allowances the agent will infer it has.

I think that’s the real tradeoff. The metaphor brings in behavior we never wrote down, which is both a part of why it works as well as a risk. One reviewer agent declined to send a PR back over a docstring nit “while the patient is exhausted.” It was a fair call, but there's no explicit instructions in any of the prompts/skills that said to do that.

We contain it by keeping the persona to a few sentences and writing the rest as explicit scripts or workflows. The rules that really have to hold aren’t left to the role. Human merge approval is a script checking for a real GitHub approval on the PR. A CI failure or a merge conflict bounces the PR before a reviewer agent is even launched. So the metaphor shapes judgment calls, and the hard rules about what should or shouldn't happen are in code.

I think there’s a famous book about man hours that attempts to answer this question

That’s an interesting perspective. I was viewing the challenge here as being driven by the fact that individual agents write worse quality and less reliable code than individual humans. The idea that the challenge is predominantly about integration at scale is different. I’m still unconvinced the medical analogy is necessary though.

Humans: $160k / 9 months = $600/day

AI Software Factory: $4172 / 2 days = $2086/day

This seems unsustainable, unless you're also generating 3x the revenue.

I don't get your maths? [1] 160/9 months != 600 for any given number of days a week e.g. 5,6,7 (something between 6 and 7), but where did the 9 come from, is it some kind of adjustment for weekends? Anyway there is also a concept of "fully loaded employee cost".

In any case the way to think of this is not "/day".

The reason is simple. If you buy 1000 barrels of oil, you buy 1000 barrels of oil not 17.4 days of oil. There is now a disconnect between work done and time. Infact you would be sane if you said "that result I can get in 2 days for $4172, if you can get me that same result in 1 hour, I'd pay $8344". See where this is going?

Yes a lot of thought work is now a commodity, and if you want the commodity faster (last minute booking, uber to come quicker etc.) you pay more not less. Value being $/hour is over.

[1] The ? acts as both a question and a regex.

Its an approximate 160k ÷ (9 months x 28 days a month) = 635.

I dont think this is comparable to buying a barrel of oil at all. A software factory like this, I'm assuming is going to be run 24/7.

I get what you're saying, but thats assuming someone's buying their software at the faster pace as well, which we don't know.

I have qualms and disagreements with the whole "software factory" concept, but your math doesn't help the argument. The cost of human labor is more than salary—or even hourly rate, if you're thinking of paying a contractor.

I'm comparing cost per $ of revenue, not labor. And my math is for a scenario assuming a company is running this type of software factory year-round.

Now one could say that the same feature that a customer would've paid $1M for was done in $4k instead of $160k. But buyers are going to demand much more as they see this kind of pace. You'd be able to charge maybe only $10k, because your opportunity cost goes down since you only spent 2 days; you can't keep the same margins that you needed to in order to justify spending 9 months on it. I think it gets very complicated and we're yet to see these economic effects.

Only one of the two trends is downwards.

... Token cost is still heavily subsidized by the frontier labs while they're private companies and raise billions whenever they decide. The downward cost in tokens is unsustainable long term without revolutions training or inference of models.

There is no evidence whatsoever token cost is heavily subsidized.

I run Deepseek v4.1 flash and pay per token.

I get more intelligence and mileage out of few tens of euros on their APIs than I would've got out of Codex+Claude 200$ subscription in June.

If you don't see the downward trend that's a you problem.

[flagged]

Reminds me of the “surgical team” development model in the Mythical Man Month.

I thought this too. It's fun that they were working with DB2 as well.

Nice writeup. Structured handoffs and external plan reviews are good takeaways.

Thanks! And thanks for reading.

[deleted]

I absolutely love how you are able to pull so much latent behavior from the underlying LLM. I wonder what other analogies can be pulled into agentic coding that come baked into the existing weights.

The patient “leaves” when the bug is fixed?

What would be the analogy if the patient dies?

The closest thing we have is if a merged change has to be reverted, the Safety Department holds a Morbidity & Mortality conference and writes up what went wrong. A patient that needs three or more rounds of rework gets one too. There have been seven reverts so far, so not many funerals.

Bug becomes a feature

This is fantastic work. I've been exploring my own work system here: https://github.com/mas-bandwidth/nova-sprint and I'm adopting your ideas. Thanks!

[flagged]

congenially my agents started calling bugs gaps

That's hilarious, tell us more!

Jesus Christ, we are just clowns aren’t we.

"Just lost another bug."

"It never gets any easier, huh?"

It brings a whole new definition to some medical terms like the "crash cart" used when a patient "codes" [0]

[0]: https://youtu.be/3gRazPZbYOo?t=60

[flagged]

[flagged]

[flagged]

[flagged]

[dead]

[dead]

[flagged]

[flagged]