That makes sense to me as a tactic to prevent that problem.

But the article talks about how so much code was written so fast. Seems to me that to create that much code that fast you have to have AI produce both the code and the tests.

So I am not sure in this project if the AI can edit the tests or not. I assume that because this is Anthropic the AI is doing as much as possible, which would include editing tests.

Bun had an enormous existing test suite written in TypeScript. Getting those tests to pass against the Rust version was the key thing that enabled the project.

You can go and check if the AI edited the tests yourself: look at the git history of those files in the public Bun repository.

But Bun also implements the interpreter to those tests— doesn't that mean it has a lot of leeway to over-fit to them? It's not as clean as "don't change these files" in this case.