The AI labs already figured it out:

1) reverse engineer the code 2) train a model on the code 3) use the model to write the clean code

Step 2 is the key “cleaning” process

So maybe a good strategy would be to use something like REA, put it on GitHub, wait for the LLMs to train on it, then just use the frontier models

/s

I prefer the word *laundering

LLM's have already trained on similar enough code

Even better, 1 and 2 done, just move to 3 and profit