You don't need documentation or the 3rd party memory systems. The code IS the documentation.

All this stuff is LLM rube goldberg machines. It just pollutes context.

I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.

I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.

Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.

But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.

Rejoice! The LLM will write the docs for you, relieving you of most of the work.

However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.

Yeah. Basically IMO/IME the current best practice is to start the LLMs out with minimal/no skills/instructions/docs.

Notice the things they struggle with and the small problems (typically, environment issues IME) they repeatedly encounter and re-solve across multiple sessions. That is what your instructions should cover. When possible, move those instructions into skills, so they get loaded into context selectively instead of on every session. (Example: instructions for running specs, placed into a skill that only gets loaded into context when it’s time to run specs)

Again, this can be automated by the LLMs themselves: both Codex and Claude (and I’m assuming other major harnesses) know how to read their own transcripts and are good at looking for repeated friction and making concrete suggestions to reduce that friction in the future.

It takes a bit of a time investment on the user’s part, and every now and then you probably should throw it all out and start fresh so that the new batch of instructions can be appropriate for the current state of the repo and the capabilities of whatever model(s) you’re using.

100% agree. Most of the things that make development better for humans also make development better for agents… and I think docs are even more important with agents, because of some kind of multiplicative effect. The agents are coding faster, and the benefits of documentation are somewhat more pronounced because of the speed.

Wow its utterly impressive that people dont understand this. I create docs and maps to instruct my LLM and im always keeping my docs up to date..

In general, as LLMs get smarter, “less is more” becomes increasingly true. You’re polluting their context with dozens or hundreds of instructions, all of which the LLM tries to satisfy. However,

    You don’t need documentation […or…] memory
…oh heck no. Easy to miss at Claude and Codex’s default detail level but if you read your actual session transcripts, you are almost certain to notice the LLM solving lots and lots of the same little problems over and over again.

    The code IS the documentation
First, lots of things cannot be learned from the code.

Trivial example: LLMs struggled with QA on our app. They didn’t know how to find the seeded test accounts. They would create new ones and do it wrong, or find the seeded test accounts but not know their passwords because they were encrypted, so they’d change the passwords but not tell the other agents. Shitloads of tokens burned. It was a no-brainer to just give them the credentials in a skill that gets loaded when they do QA.

You could say that’s an environment issue, not a code issue. But as far as actual code IME at a minimum we have to tell the agents our general repo structure and architecture patterns.

The days of “you are the world’s greatest Python programmer, write good code” or whatever are over (if that style of prompting even worked in the first place) and as I said less is more. But, still….

> The code IS the documentation

I liked this advice when humans wrote code. Though even then I'd urge people to write meaningful commit messages that capture the "why" of what they did, so no one tramples their intent by mistake.

But not sure it works in an age where most code is LLM-generated. Especially if that code is not even reviewed by humans (irresponsible or not, it's happening), and commit messages are also generated by AI. I think something is needed to separate "what did the human operator intend" from what the agent went and built.

I do agree that this gets way overengineered. My approach has been more or less what you stopped doing though - committing all our timestamped "plan/implementation docs" and "investigation docs" that document what the user wanted + empirical findings, and making all prior session transcripts searchable. It's seemed mostly helpful? For whatever reason I haven't run into many staleness problems so far.

Code as documentation works better when the code is declarative or a DSL. These capture intent and promises (as in promise theory) better.

When it is not, it has to be reasoned out and does not work well for documentation.

Other things that code and tests alone do not capture well:

- promises (as in Promise Theory) made to other parties. Claude already has PT in its training data.

- Constraints-inducing-properites, as in Roy Fielding / Christopher Alexander. While tests, and property testing can capture properties, there is no formal connection to the constraints that induces fhem. By constraints, I am not talking about business requirements and business value — those are better understood through Promise Theory. I am talking about things like at-least-once delivery or total ordering (from append-only constraint). Claude already has Fielding’s dissertation and Alexander’s works and ideas in its training data.

- grammars, as in pattern panguages (not just patterns) a la Alexander / Fielding are also not captured in code alone. These tell both humans ans AI how to extend a pattern, and how to identify anti-patterns (when they violate a constraint-inducing-property)

- LLMs are trained with many different worldviews and bounded contexts at the same time, and is very capable of translating across it. However, these need to be soelled out, otherwise it would talk in whatever it infers

Specifications written for the exact way components are wited together run into that stale doc problem. Although it takes much more human attention and token burn to describe things in terms of pattern language and promise theory, it becomes easier over time. The actual implementation plan tends to fall out more cleanly when all those other stuff are at least considered. This is where I have been spending most of my time.

It’s a weird thing isn’t it, the urge to save these artifacts? The worst is when the llm refers to the decision and design in code comments. In my opinion there’s one thing that is worth documenting; tricky architecture or implementation details that are some how counterintuitive to what would have normally been done. But again this can be documented in the code and tests.

Code typically documents the "what" and "how", not the "why".

Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.

And agents are pretty bad at inferring when to do, and often fall back to verbose clutterin the comments.

[flagged]

I don't understand. Code doesn't capture the context in which decisions were taken: why is code the way it is? What is important? What is not? How can agents make correct decisions without knowing context that cannot be inferred from code?

It's frequent for SWEs to make blanket statements with considering the vast space of issues other people face that they don't have experience with or awareness of. Anyway, I'm not going to tell you what you do or don't need, only what worked and didn't for me.

In my codebase it is difficult to get agreement on comments and documentation so rather than rely on it I adapted. One of the first things I did when I succumbed to agentic development was to point codex at the code and ask it to generate a high level description of where important files, such as our public API, reside, what the hierarchy is, what the code does, etc. In my case, this level of documentation is fairly static if I avoid implementation details. So now I have a handful of agent files in my tree and it seems to save quite a few tokens and improve my results. I frequently have other devs ask me how I get such good results when doing agentic reviews of their changes(always my first step now before I start my human review). I also include instructions in the agents files instructing the agent to maintain the agent files if any relevant changes are made. It seems to work quite well for me.

I partially agree, way too many people are cargo culting overly complex AI workflows with little empirical data. My framing is a bit different though, I consider code the spec and actually keep a decent amount of docs for higher level concepts. So far this is working well for me across Claud and Codex.

Idk I just can't agree. Code is the what but it doesn't tell you the why. There are so many times where at first glance the code seems suboptimal or bad or wrong, and it's only when you learn of some constraint somewhere else that it begins to make sense.

All code is written under constraints, and most constraints live outside the code.

I went through the same cycle as well.

I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.

What about big projects, where much of the code is not written yet?

Code can get much larger than the documentation that summarizes it. It also doesn't cover intent or rationale. Even if you inexplicable don't want documentation, at the very least use something like gitnexus to map out your code because relying on code alone isn't good enough

Yup. Examples examples examples. All of the descriptive stuff is just nonsense that confuses the point. Makes perfect sense when you remember that these things are not intelligent, but truly just autocomplete on steroids.

>> You don't need documentation or the 3rd party memory systems. The code IS the documentation.

We have heard this nonsense from the "we don't need to write comments, code should be self-documenting" types for decades. It was wrong in that context, and it is wrong in this one.

Code tells you how a system works. It does not tell you why it works that way. That is what memory is for. It exists so that your AI does not keep undoing past decisions when it writes or refactors code.

[flagged]

[flagged]

[flagged]