I think programmers are frustrated in part because the code used to be our place to do our thinking, but now LLMs just buzz through the file changing thousands of lines and we can't keep up.

It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that. We just have to raw-dog it by thinking really really hard and remembering how all the code connects together.

I want a tool that gives programmers a place to record their thoughts. Developers need a place to draw and write, and also interleave blocks of code that automatically update to match the actual state of the code.

The closest thing I know to this is org-babel, part of Emacs, which allows you to push code blocks out of an org file into an actual source files, or pull them in from actual source files. This is mostly done manually by invoking functions called `tangle` and `detangle`.

I intend to investigate this further in Emacs, since I'm an Emacs user, but Emacs is never going to be the friendly UI we need to make this tooling common.

> It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that.

Doesn't the LLM do that, if you want it to? I do use it to write code, but the more striking ability it grants me is a means of understanding legacy code far faster. I can ask "what actually causes this branch to be taken" and it's usually right. Ok sometimes it's not right, but I'm not always right either even after I spend tens of minutes reading code.

It's also exceptionally good (i.e. fast) at looking through git history to figure out where/when a certain behavior originated, which can be difficult (time consuming) if code is continuously being refactored.

Sure you could use it to vibecode. I don't, rather the opposite, I understand my own changes better. But I still fear that I'm going to be obsolete as soon as it figures out what questions to ask. And I'm unwittingly training it to do that.

I feel like we also need better non-LLM driven ways (like better static analysis tools) to analyze LLM-driven changes. The way changes are presented in modern IDEs was designed around reviewing human-created changes and doesn’t really feel like it’s keeping up with presenting and validating what modern LLMs are doing.

I've been playing with often asking Claude to generate visualizations for me of what it's doing in the sense of diagrams of different types. Sometimes it's as simple as "show me what you're doing using an infographic". Othertimes I'll ask it for sequence or class diagrams or bipartite graphs if it's designing some sort of mapping. Bipartite graph is also very helpful for following the plan of a big branch rewrite or squash which I often do before making a PR to get rid of all those confusing in-between commits.

I find forcing it to visualize things immensely helpful. I'm usually studying git diffs but when working of a big feature or refactor that can just be too hard.

I've never been very pro "visual programming" and always hated UML et al, but part of me is starting to wonder if it's time for us to give it another serious go.

> now LLMs just buzz through the file changing thousands of lines and we can't keep up. It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that.

I do it differently, I focus on better recording what the user wanted, the so-called "user intent". To do this, I record all messages typed by the user since the start of the project, whether 3,000 or 10,000 messages. An LLM can churn through them in 10 minutes and derive a fresh, up-to-date interpretation from the raw data. This can be used to judge whether the implementation has diverged from the intent, or, in other words, to realign the code and tests. The messages the user writes are usually designs or corrections, a very rich, compact signal. If the user struggles with something, it could result in a tool, a skill, updates to the project docs, or new tests.

But a codebase is a state machine, and given the somewhat random style of LLM outputs, those amplify to an extent that the recorded intent needs to include the outputs too, in a way? Or are you basically doing high detail specs?

Yeah, "Intent-based UX" is evolving to address this challenge, but it's got some catching up to do AND is far from mainstream....

Yeah, this is the tough part. In order to write code that worked, you had to have some kind of mental model of it. Now that's not true.

Now, when someone sends a working PR in, even high quality and well tested, they may actually have no idea how it works.

Which, in my experience, means they sometimes cannot fix the bugs they've introduced.

> It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that. We just have to raw-dog it by thinking really really hard and remembering how all the code connects together.

I deeply relate to this. When engineers were writing all the code that meant every part was deeply understood by _someone_ on the team, and they could valuably contribute to maintenance and further development. It wasn’t perfect, people leave, people forget things, etc, but the overall coverage was high and valuable.

Now, every agent-produced MR introduces code that is deeply understood by _no one_. It’s the “original developer left five years ago” problem, but now growing on every single new piece of code. Reviewing doesn’t give you the same depth of understanding, and the continually increasing impulse is to just approve, maybe nudge it about some isolated enum types or something, but don’t take the time to understand it, just keep the train going.

But then what happens when something breaks and the cloud agents are down…

I've been having agents build knowledge graphs, they're a tremendous mess to start with, but I take the time to manually drag nodes around or group them in meaningful ways so that it's actually human-browsable. This is boring enough to create space for me to think in. It leads me to go on expeditions into the code which surface the missing details. It's also a nice way to communicate context to agents. Like, I can hide all but the relevant nodes from an agent before suggesting that it query the graph to understand which service references which other service via which api, which database tables are read/written by such an action, etc...

The annoyance I have is that it isn't one of the other. I spend a lot of time thinking and fighting with whatever frontier model we're on today. I've seen people let LLMs make absolutely atrocious decisions and just click next and collect checks.

I just posted this morning about this type of engineering has secured us numerous customers and put some projects in our backlog that either need significant rework or at least a very close eye to see if their issues crop up.

IDK, the idea that we shouldn't be thinking is a worrying one. I still have to think a lot.

What I'm _mostly_ worried about is that the path to get to high performing senior is basically a burned bridge with our current training techniques, and I'm not sure we'll adapt before a brain-drain situation in the industry.

> It's still valuable to deeply understand parts of a program

Part of woe is that once you've reviewed, validated, and comprehended a piece... Later gets casually mangled by some other LLM-generated urgent change.

> I want a tool that gives programmers a place to record their thoughts. Developers need a place to draw and write, and also interleave blocks of code that automatically update to match the actual state of the code.

Sounds, like you already mention with org-mode or similar ones) like literate programming (https://en.wikipedia.org/wiki/Literate_programming) or jupyter notebook.

I think the solution is still code, just at a much higher level of abstraction. Maybe a start is kind of typed ADR or FSM that guides (constrains) the agents. I believe more type checking guarantees will be more and more important for agents.

I find that there are domains where LLMs are much faster and more skilled than I am, particularly in extremely well-documented but technical and complicated, but a lot of domains where they cannot do anything at all (mostly novel issues, weird architecture issue resolution, etc.).

Building a basic X11 window manager is almost a one shot prompt.

Modifying a UI toolkit to make it work with MSAA/IA2 is simply not possible.

There's a lot of room for deep work left... for now.

When you say this is simply not possible you sparked my interest. I would be curious to chat about your approach?

If I were trying to accomplish this particular goal I would first consider what the agent could see. In particular does it have an accessibility inspector of some kind? or even NVDA hooked up with NVDA Remote so that it can actually see the implicit a11y tree for the toolkit it is working on? My email is in my profile and I would love to chat about this.

The workflow I use to still keep up with everything is to start coding by hand and only once I have a good idea how the rest is gonna look like and am bored I had off the rest of the pr to the LLM.

Could be just defining the methods without filling them but depending on the mood I code more by hand or less.

Yeah, I can relate to this. However, I haven't found it too difficult to adjust. I have found myself creating draft PRs, and then just sitting on them and thinking about it for a day or two before I even consider merging it. This lag time is the time that I used to spend typing it out and thinking as I went along, now that's happening later. I have closed more than a couple of my own PRs once I had time to consider them. I am rarely shocked when I wake up the next day, look at it, and think: "Eh, this change is not sufficient because it doesn't address X".

> It's still valuable to deeply understand parts of a program, but we don't have any tooling that helps us do that. We just have to raw-dog it by thinking really really hard and remembering how all the code connects together

Which is frankly exhausting to do when you have to keep up with the rate of LLM changes

I can sling code at about 20x speed with an LLM, but I can only understand it well enough to support it at 5x speed and I can only make decisions that won't piss off the rest of the company at about 3x speed. My job as a software engineer is to therefore slow down to working merely 3x faster than before despite the extra headroom that the LLM gives me. Anybody can give into the seduction of new features poorly understood, to be a specialist means to bother spending the extra time.

Or at least that's the current model I'm playing with.

This is my dilemma as well currently, because there's no obvious place to draw that line. The boundaries are all subjectively defined.

I could change a whole UI completely in 30 minutes to something fundamentally better but then 30 people would all wake up and be upset they weren't consulted and need training for it. That training and consultation will take hours and hours. And probably generate feedback - some of it correct, some of it misguided - that needs to be human negotiated, taking more hours. The effective maximum rate of change is limited so dramatically more by other factors than the technical implementation that we have to completely redesign process now to cater to those factors.

We are in a weird space now because most of the process is still built around a presumption that technical implementation is a lot of work. The main reason to be upset that you weren't consulted about a change is because there's a presumption that you will be stuck with it - ie: it's a lot of work to change it back. But it isn't a lot of work, it's effectively free. All this is just living in inertia right now.

I've been handling it as a sort of voluntary A/B test.

A is what you're used to, B is what I recommend. If I can convince people to start using B instead, I can look at the metrics for A and conclude that it's effectively dead, and then I can remove it.

It's working out for me, but maybe not a fair comparison because I only have something like 15 users.