tbh I thought this was a commonly used technique even prior to LLMs? I know I've been using it extensively myself, but I was inspired by Dolphin's extensive visual CI system.
tbh I thought this was a commonly used technique even prior to LLMs? I know I've been using it extensively myself, but I was inspired by Dolphin's extensive visual CI system.
If it is common in the world of video game porting, that just shows my ignorance. I'm familiar with visual diffs in CI for e.g. web development (comparing a static component), but to do that to compare frames over time in a video game/3D environment is new to me.
There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.
Even then, visual diffs were pretty flakey for web development, because one's OS and browser choice would slightly alter the exact pixels blitted to the screen. At least this was the case for the tests that would simply match pixels instead of computing a sort of visual hash.
It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).
For web UI tests this is mainly solved, at least when using Playwright. It allows setting thresholds, percentages and some other config items to allow some small differences in pixels. https://playwright.dev/docs/test-snapshots#options
I wouldn't call that solved, no.
That's an attempt at mitigation, but most definitely not solved.
It still causes both false positives and false negatives through that. The only true "solution" is to make sure generation always happens on the same platform as your ci... And various mitigation strategies around that (eg fall back to structure tests on other platforms vs actual visual diffs in ci
Also not unique to playwright. Been available basically everywhere since the start
I'm getting a 1997 PC game to run on modern hardware and fixing bugs and upgrading graphics as I go, and the amount of quality support tooling Claude is producing along the way is impressive. Fully headless in-memory execution (which, among other things, is used by it for per-pixel diffs too), logic VM devompiler and visualizer, asset explorer, CRT simulator... I just say what I'd like to see, and Claude does 120% job on it each time.
I do that on my renderer. each commit when it's ready to be merged gets a class of visual diffs. Dolphin's way was a major inspiration how to structure it, but it was _the way_ in rendering way before it.