Claude ran npm update (update all dependencies to the latest version compatible with the semver specified) in my repo without telling me when trying to fix some problems. Given that only updates the dependency lock-file I didn't notice and it caused several hours of debugging for me.
It is quite sneaky how LLM output can sometimes bypass human verification like that. No one is going around checking every single line change in auto-generated files. Someone could easily sneak a malicious dependency in there through some online tutorial that the LLM searches for.
> No one is going around checking every single line change in auto-generated files.
There's a simple fix for your particular case: commit your lock fine (which you should do) and always review the diff (which you should also do). (:
You've made me realise a good signal for bug hunting: Search repos with lock files listed in their .gitignore.
It's the sort of terrible practice that someone might be frustrated into taking after a nasty merge conflict, and signals a willingness to cut corners.
My lockfile was not gitignored, I had made significant changes to package.json so I was expecting diffs in the lockfile.
I just don't usually read lockfile diffs and claude inadvertently updated a few dozen packages to new minor versions without me noticing. In fact I only realized the problem after I looked at the lockfile diff.
Makes sense. This is why I like using jj, because "update deps" and "fix bug" are going to be two separate commits, and when "fix bug" has changes in the lock file, the red flags go up.
Claude seems to be super good at jj so that can take the edge off as well.
Why do you need jj for this? I'd usually make incremental commits with git, and I'm wondering how jj levels it up.
> You've made me realise a good signal for bug hunting: Search repos with lock files listed in their .gitignore.
What would be the point of that? Do you just go around hunting for bugs in random repos?
Sure, in the spirit of open source, why not? It's a hobby, and it scratches an itch. I very much enjoy deconstructing things more than putting them together.
We also live in a world where a package written by someone learning to code ended up critically underpinning the entire ecosystem and is downloaded 500 million times a month.
Ignoring the eco-terror aspect of that for now, it means there's an awful lot of code out there which is finding itself under constant attack by a fleet of hostile AI.
I don't personally believe that the solution to that is "more AI", which firstly just overwhelms maintainers and secondly surrenders our human agency to a giant machine, with a hope that the "good" side can out-spend the bad.
Nor do I think the solution is to abandon the open internet and retreat behind corporate walls into curated spaces, "benevolently" protected by giant companies.
Which means holding on to the open internet requires a human approach, and any signal to help amplify the work there is a benefit.
>We also live in a world where a package written by someone learning to code ended up critically underpinning the entire ecosystem and is downloaded 500 million times a month
whoa what? which one is that?
As another commenter said, it's "is-even":
https://github.com/i-voted-for-trump/is-even
From that page:
> I created this in 2014, when I was learning how to program.
I've nothing against Jon Schlinkert, it's not his fault the way we build software is more than messed up, where our build systems are so brittle that, "Throw out the universe and rebuild it from scratch" became not just acceptable, but the main way to get build systems to work reliably.
Check is-even and is-odd npm packages. https://www.npmjs.com/package/is-even
That's still quite a ways away from 500M+ downloads a month, more like ~4M downloads a month.
Still a huge number of downloads, don't get me wrong!
You're right, I was reading the stats for "is-number" and mixing them up for "is-even":
https://www.npmjs.com/package/is-number
170M downloads / week.
Same author, similar vintage. Arguably a necessary package, but that just further indicates how messed up javascript was.
So nuts.
> Arguably a necessary package
Arguably a somewhat important part of a standard library!
A large part of the problems of Javascript are corollaries of lacking of a good standard library, and the relatively long time it took and is still taking to fix that.
It wasn't really until ES2015 that a better standard library really started to take shape, and, thanks to IE11, it was a very long time before that didn't need poly-filling.
In a sane world, you'd just parse whatever you're after and then check for NaN or null.
You can't do that. Pop open your favourite javascript runtime and type:
It's also a good practice when taking a new job, especially if someone is a contractor and changes gigs every few months or years.
I've found that, when I start a job, I have to rely on smells like this to know what kind of mess (or if there is a mess) I need to clean up.
That's a manual step, not a solution. the solution is just boring basic file permissions. Treat Claude as semi hostile user. If you don't want them accessing your files, set the permissions to exclude them (like require sudo).
I already do this for my unit tests, because Claude will "fix" the tests so they'll pass.
I made changes to my dependency lists in the same code where Claude ran npm update. The lockfile diff was a few hundred lines after I undid what Claude did.
And yes, eventually I did check the lockfile changes and spotted the problem. I just usually don't check the lockfile that throughly.
> I made changes to my dependency lists in the same code where Claude ran npm update.
...but was it in the same commit? Two "update lockfile" commits, one yours and one Claude's should have made this obvious, no?
Here's another useful rule of thumb: never mix your changes with the agent's changes. Agent always starts with a clean repository (no pending, uncommited human changes). You always start with with a clean repository (no pending, uncommited agent changes).
Personally I have this in my `AGENTS.md`:
So my workflow is usually this: start agent with a clean repository, tell it to do a thing, it works in the background, then once it's finished I come back, review, rewrite and clean up half of what it wrote, then maybe iterate some more with it, and finally do an interactive git rebase to get a clean commit history.I don't let claude make commits for me, every commit is my own. I check the diffs before commiting.
> never mix your changes with the agent's changes
yeah, but if you don't want to lose your existing context sometimes you have to. When I do, I tell claude to check the diff on the files I changed, which is quite annoying to be honest. But still easier than telling claude to do _very specific line-change_ on file X.
> yeah, but if you don't want to lose your existing context sometimes you have to
Sorry, I'm not sure I follow. What do you mean by "lose your existing context"? Can't you just... commit in turns? It's not like you're editing files while your agent's also editing in parallel, right?
Again, the trick is to treat the commits as throwaway checkpoints/packets of work. They don't need to be pretty, nor need to make sense. The point where you clean that mess up is when you're done and you're doing an interactive rebase at the end. At least that's how I work.
Ah okay, I thought you meant to fully close the session before you make any manual changes.
In my company we use graphite and stacked PRs so it is highly encouraged to keep one commit per PR, so I am constantly ammending my commits.
> In my company we use graphite and stacked PRs so it is highly encouraged to keep one commit per PR, so I am constantly ammending my commits.
So this makes it even simpler for you. Then you don't have to care at all about keeping your commits clean (as in: you don't have to keep them organized enough to be able to reshuffle them into a nice set of multiple commits later on).
Just commit whatever, and then just do `git rebase -i` interactive rebase at the end to squash them. You don't have to keep amending the same commit over and over again!
Heh, I also notice some coding agents like to explicitly git ignore the lockfile.
“But I have an agent for that.”
I cannot tell you how much time I have saved by stopping Claude and asking, "what are you doing?"
At least 50% of the time, Claude "realizes" it already has all the information but is doing something that's unnecessary for the current work, stop, and tell me the previous step has completed.
People complain about approval prompts etc and have Claude run in fully autonomous mode. Outside small bug fixes, I just never find that useful. It helps me immensely to see what commands Claude is running to understand where the work is going.
Also Claude often cooks up atrocious overengineered ideas but responds pretty well to being guided hands on to the desirable scope.
When testing Claude code in auto mode in a fresh sandbox with a docusaurus website freshly cloned, I asked:
Can you see the docs folder with the git project?
It was in auto mode. So it immediately saw a docusaurus site (good) but instead of stopping there, it installed nodejs from a static binary download (no root access so only way), ran npm install, started the dev server and confirmed the project worked.
That is some crazy amount of leeway for an intent based classifier. I'm not surprised it's full of holes, and it seems to be 100% by-design.
Worse: Claude installed packages by just typing versions into package.json instead of running `pnpm install x`, then when running `pnpm install`, discovering that the package versions are too new and incompatible due to the default `minimumReleaseAge`, then proceeding to circumvent this by disabling `minimumReleaseAge` and running a full package update :)
Maybe I don't understand you correctly but if your lock file isn't in Git then you have bigger security issues than LLM output, given the last years NPM worms. Unless you're a single developer and the file on disk is the primary source of truth.
[dead]