Oh my god, and the author even supplied a "proof"[0] visual diff harness... that it replicates the original game pixel for pixel.

Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.

[0] https://github.com/terrapapagalli1516/quake-srp/tree/main/or...

tbh I thought this was a commonly used technique even prior to LLMs? I know I've been using it extensively myself, but I was inspired by Dolphin's extensive visual CI system.

If it is common in the world of video game porting, that just shows my ignorance. I'm familiar with visual diffs in CI for e.g. web development (comparing a static component), but to do that to compare frames over time in a video game/3D environment is new to me.

There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.

Even then, visual diffs were pretty flakey for web development, because one's OS and browser choice would slightly alter the exact pixels blitted to the screen. At least this was the case for the tests that would simply match pixels instead of computing a sort of visual hash.

It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).

For web UI tests this is mainly solved, at least when using Playwright. It allows setting thresholds, percentages and some other config items to allow some small differences in pixels. https://playwright.dev/docs/test-snapshots#options

I wouldn't call that solved, no.

That's an attempt at mitigation, but most definitely not solved.

It still causes both false positives and false negatives through that. The only true "solution" is to make sure generation always happens on the same platform as your ci... And various mitigation strategies around that (eg fall back to structure tests on other platforms vs actual visual diffs in ci

Also not unique to playwright. Been available basically everywhere since the start

I'm getting a 1997 PC game to run on modern hardware and fixing bugs and upgrading graphics as I go, and the amount of quality support tooling Claude is producing along the way is impressive. Fully headless in-memory execution (which, among other things, is used by it for per-pixel diffs too), logic VM devompiler and visualizer, asset explorer, CRT simulator... I just say what I'd like to see, and Claude does 120% job on it each time.

I do that on my renderer. each commit when it's ready to be merged gets a class of visual diffs. Dolphin's way was a major inspiration how to structure it, but it was _the way_ in rendering way before it.

The visual diff harness was probably how they got rid of a lot of visual bugs, just tell the LLM to keep going until the pixels match exactly as the verification criteria

Using pixel data might be unusual, but anyone porting a game with a replay feature is going to realize it's an easy way to compare implementations.

There are several good algos used for diffing graphics for not so pixel-for-pixel rendering allowing thresholds (you are going to end up with differences between software and GPU rendering at least - it's inevitable):

https://github.com/NVlabs/flip

But also SSIM, perceptualdiff, and many others

For UI/web certain other methods are better, for video - I'm sure there are prefferences there too.

I suspect it's inspired by gbaeval. Both have the "oracle" and other similarities. https://gbaeval.com/

I've recently gotten an Oracle in my translation system too, I think it's a common need to compare bytes betwitx to different systems. I didn't think it was specific to one company.

Yeah cool.

How long before the same thing is done to like, banking back ends? Wallstreet proprietary software? Amazons logistics and distribution systems?

It seems like we might be weeks/days/hours before a situation where someone back engineers and spoofs a system so pivotal to modern human society that the plug needs to be pulled.

Like I've said many times: LLMs are useful for this (i.e. porting software from one programming language to another).

Porting software is painstaking grunt work which still takes a moderate amount of intelligence. It's therefore extremely expensive to port say, COBOL banking software running on mainframes, to another language like Java. That's why a lot of COBOL software is still in use. I expect this to die out in the coming years as many of these systems will finally be ported to another language (could be Rust or any other language).

In many ways the financial system is a distributed monolith. COBOL is still around because of bugs and issues (sometimes in code, sometimes in the OS it ran on, sometimes in the computer hardware) that have been around for decades are often dealt with downstream with their own exceptions. This is why IBM still sells mainframes. They literally emulate software and hardware bugs from equipment as far back as the 1960s.

I spoke with a semi-retired COBOL programer about a decade ago. They literally often can't fix specific bugs because multiple other consumers (from banks, hedge funds, insurance companies, the central bank) outside of their organization (and subsequent downstream consumers as well) would all need to adjust their code. Rewriting it is grand, but won't fix the spaghetti code mess. This is made even worse in the US, which has a much more fragmented banking system. "Modern" COBOL isn't even that bad, but they can't even use it most of the time. There is SO MUCH that isn't even documented - and if it is, is in a binder that's been shelved decades ago and half of this developers time was going into the corporate archives to dig it up.

This isn't to say that a lot of legacy code can't be modernized, but it's not always that easy.

If someone recreates Amazon's logistics and distribution systems they could try to compete with Amazon? But they'd also need the connections, distributors, transportation, etc. same with banking software, you need capital to be a bank not just software, and if they have the capital then the technology is working we intended making it easier to make new things and innovate, or at least just compete?

No, I am not talking about "taking over" companies and trying to emulate them and do business yourself. You just need to be able to break trust in the api calls and no one knows if a purchase order or transaction is legitimate.

Obviously you need to have access to the keys, BUT I don't see this as a dealbreaker anymore because you just get your agents to go and find them.

I think you’re saying ‘being able to do this means the opportunity for more fraud, by producing fake XYZ as proof’.

Photoshop has been around for around 35 years, fraud has always been an issue. There are plenty of reports of people selling things via Facebook marketplace and the ‘buyer’ showing them sending a payment on a fake baking app. Fraud will always exist and I don’t think tech will make it worse, everyone needs to be more cautious and tells friends and family to be the same.

Nah small potato stuff.

I'm talking about cloning hmi platforms to send fake instructions to offshore oil platform valve bodies or insert false market trades to collapse companies.

Ah. So in that instance are those platforms not validating that the things submitting information are correct and true. A bit like a utility company needing to do manual reads every now and then to ensure they are getting a correct signal.

Or perhaps uploading faulty firmware to centrifuge controllers?

Or they can just take your money and not send out anything.

Alternatively, you just act as a middleman drop shipper and slightly raise the price more than Amazon’s and skim the difference. It might be a while before they find out.

Example, parking places with QR codes for paying webapps.

Currently a plague in some European countries.

It looks like the real site, and you pay twice, in the fake app, and later the police.

That’s an insecure design. The way we do it here is that you install an app and register your register number and payment card in it. Then when you drive in and out from the parking lot your license plate is scanned and you’re automatically charged. There’s only two providers so it’s not a huge hassle, if there was a single app per garage it would not really work from UX perspective.

No room for hostile social engineering.

Are you sure to install the right app though?

https://www.bbc.com/news/articles/cwyjqg578e1o

And given your German nickname, here isn't safe either in a general way, when folks aren't regularly parking on the same place.

https://www.adac.de/news/verkehr-quishing-parkautomaten

Yep. On a small scale, you could skim money off transactions. On a large scale, you could break global distribution and logistics chains.

> spoofs a system so pivotal to modern human society that the plug needs to be pulled.

why would a recreated system be detrimental?

If currently there's a monopoly on a software, this AI recreation is a good outcome to poke holes in that monopoly. It's only bad if you are financially invested in said monopoly, and this would be a minority compared to the amount of benefits that society at large could obtain.

This is basically a digital era anarchist view - the problem you're overlooking is that a lot of critical infrastructure we rely on runs on systems that are considered security through obscurity. Software most people probably wouldn't even know or care that it exists. If you can break the trust of vendors by being able to spoof their proprietary platforms, a lot of the highly efficient networked systems becomes vulnerable to injection and abuse if you can't trust whos making calls to it.

In the past you'd need nation state actors with considerable budgets to do this kind of thing, and we're on a trajectory that could see any kid in his bedroom could do it.

revealing that security thru obscurity is broken can only lead to a better future, even if in the intermediate one there are lots of breakages. It's suffering that needs to happen, and better sooner rather than later imho.

And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded. Like using a photoshop replacement.

> And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded.

I am, actually, talking about MITM attacks that are much further in scope than just defrauding some people using their banking app. I work in resources and operate HMI systems that are networked, but not exactly the bleeding edge of modern software development. If you had an ability to decompile it and recompile your own version you could start sending instructions to infrastructure all over the country - the only thing stopping you is the keys, which if you're intent on hacking someone you'd have the means to obtain anyway.

I can see a lot broader attack vectors than just stealing peoples money. It's the erosion of trust in the api calls themselves.

> It's the erosion of trust in the api calls themselves.

and that's why security thru obscurity fails. It just hasn't so far.

And if this is the current state of affairs, then the change and pain is what needs to happen for the system to improve.

For example, all api calls would have some form of authentication and attestation.

>banking back ends

As someone who works in a bank: depending on the exacty subsystem of a bank the answer is from "already" to "in 3-5 years".

Making predictions for in 5 years with the current dynamic of the ecosystem seems rather speculative

[deleted]

I’m already seeing videos of people who have used LMs to reverse engineer and clean room reimplement entire video games. I estimate this shit is minutes away from being shut down, because as we’ve all seen companies stealing is OK, but individuals stealing is heinous and a crime.

I mean some people (me included) have been begging society to pull that plug since over 10 years now.

The plug being "the cloud" and "hooking everything up to the same internet".

These confusion attacks can only confuse people, because critical systems can exist in the same space where entertainment systems and all other categories of systems live. This was wrong even before LLMs.

It won't be long until we can ask it to replicate iOS or MacOS.

Apple needs to up its game, fast.

It won't be long until corporations will start paying AI companies to block such requests or even lobby governments to make it illegal to do.

Imagine the power of this magic superpower in the hands of business owners....

[deleted]

Adding to what you both said. I run pixel checks against live websites in a real browser, and I stopped doing full-page screenshots early on: scrollbars, lazy-loaded images and font timing shift pixels between runs like clockwork.

What works for me is diffing small stable regions, one component or one flow at a time, with a small per-pixel tolerance. I also keep a DOM-level assertion in front of the visual check, so when something fails I already know what changed structurally before I compare two images.

For the game port it's genuinely harder, the renderer decides the pixels, not the test. But the region idea transfers: lock the viewport and only diff what's stable.

You must be fun at parties

Actually Claude just did something similar for me, as I'm working on something else with Quake.

On its own it decided to do demo playbacks and take periodic snapshots, and compare them pixel by pixel if the PNGs differ.

I was also working on a web-based port but it saw the original Quake code was not modified so it compiled a native version on its own to use for this.

It had to apply a small patch to make the game completely deterministic, but it figured that out on its own by reading the code.

It then asked me to record a demo with various elements, say an explosion or being under water, and visually verified that the screenshots had those elements present.

So now I have a solid set of tests to verify against.

Opus 5.5 Medium. Used at most 1% of the weekly limit of my $20 plan.