No one can keep up with the volume of code AI produces.

We wont stop using AI.

We will use AI to check AI.

Of course this is crazy, but it will also unlock pretty insane scaling and productivity and ultimately we will manage it on either end via requirements and tests.

> it will also unlock pretty insane scaling and productivity

Insane scaling of bloat, bugs, and technical debt I'd say.

> We will manage it on either end via requirements and tests

It is so crazy that this is being touted as a sane strategy. When I was a much worse programmer, I tried to write a big complicated string manipulation function to take two types of scripts in a language and add diacritics. I had the requirements very clear. I had the tests very clearly with all the edge cases. But I didn't have a good and clear picture of how to attack the problem which was quite novel for me. As I got closer to passing all the tests it got exponentially more unruly and confusing. And nearing the end I was frantically changing little bits here and there wincing and praying and hoping the tests would pass. "Please work! Come on!" Then when I got close enough, I could never ever think about touching that mess again.

I was a below average programmer then throwing myself at some novel problem I didn't understand. Throwing LLMs that produce below average code at novel problems and relying on tests and requirements is not where we want to go to make real progress.

(Years later after much learning and coding myself I was able to redo the function in a totally different way. This time I actually understood how to attack the strange problem and made something clean, clear, and robust that just worked. The tests then become a secondary guardrail, not the main force of correction.)

We are seeing such a massive regression from what we've learned over the years of CS.

I think all code is technical debt in a way. Good code is a necessary evil, bad code is more evil than necessary.

Generating code automatically when you're not even quite sure what it is or even should be doing is insanity.

>Insane scaling of bloat, bugs, and technical debt I'd say.

You just described every legacy codebase. Many of which are widely used and do a lot of sales. You dont need a clean codebase to have a valuable product.

>It is so crazy that this is being touted as a sane strategy.

Re-read what I said. I literally called it crazy.

It is the same dynamic that gave us customer service from some call center in India. Why would companies do this? Customer service got worse. Are they stupid? No, it's just worth it. The quality goes down but the business can scale more so it doesnt matter.

AI will absolutely be good enough at doing things that we'll happily accept some jankiness at times so that we can devote an extra 3000 hours per year per person to other things.

Im not even suggesting its a good thing. I just think the incentive structure dictates it. You're not going to have time to maintain a small slice of some service by hand.

I'm not so sure LLM code today is below average. There was a time that things posted to dailywtf were normal everyday stuff

Sorry, no, they wouldn't have been WTF's if they were normal

You shared a story of a novice incompetent human programmer and this should tell us that AI is bad at coding.

It's mostly (not entirely, but mostly) finding security issues in old human-written code. It'll eventually start running out of those.

From that standpoint, it's not a crazy setup security-wise. Maybe still crazy for development.

You can point AI at any AI produced code and ask it to review it, get back 10 bullet points and a few pages of prose. And the fun part is, you can do that over and over and over again!

This happens all the time. Yesterday, I ran into an especially egregious case.

I had Fable add a new subcommand to our internal CLI tool. I reviewed and tested it locally and had to suggest several fixes that I feel like I wouldn't have had to tell a human senior engineer to do. When it finally submitted the PR, I had it on a loop waiting a few minutes for comments on the PR, then assessing/addressing/replying-to/resolving them, and then repeating again until all AI reviewers were okay with it. It ended up going through dozens of revisions and ended up with 160 comments left on the PR.

You're suggesting that LLMs get better at fixing bugs/vulnerabilities, but at the same time stop getting better at finding them? What if this difference is inherent and essential?

> You're suggesting that LLMs get better at fixing bugs/vulnerabilities, but at the same time stop getting better at finding them?

Are you implying that all code writing by LLMs atm is bug-free?

Absolutely not. By most accounts they're terrible at fixing anything other than trivial bugs in complex codebases e.g. Linux kernel, but they're much better at finding them.

So you put it in a loop and tell it to find the bugs in the code it wrote. What's the issue?

This is a self confession if I ever saw one.

Volume..... <sigh>

It used to be considered a quality of good code that there would be less code, not more.

Some people always tryin to get the highscore on golf.

You can have both less code per problem and more code overall when you make problem solving cheap enough.

In fairness at root this has been going on for awhile. No one can keep up with the volume of machine code that modern more abstracted codebases produce.

We didn't stop using syntactic programming languages we used code to check code.

Not sure it's really crazy at all. It's been an abstraction for programmers probably since we stopped soldering transistors to each other.

There is a MAJOR difference between predictable generated machine code and Russian Roulette code generator.

Of course there is.

But if you don’t actually read it…

[deleted]