Hacker News

> So we ran it head-to-head against Claude Opus 4.8: same one-shot prompt, build a 3D platformer in raw WebGL from scratch

Running a single one-shot prompt is not a benchmark, not is it representative of any sort of real-world usage.

Most agent usage is collaborative so you need to test things like reliability (when I delegate a task, does it complete it without making up test results for e.g.) and steerability (does it obey my instructions or does it just do what it thinks is best).

jameswhitford 13 hours ago [ - ]

Hi, I am the author, I completely agree! I set out to run a vibe test on this one, not a benchmark, the real benchmarks are listed. My test shows what the models can do when both tasked with a long-running, technically difficult, one-shot task.

I think your test you describe (collaborative, task delegation, task completion, TTD, steerability) is a great format for a future test that I will definitely try out.

wongarsu 12 hours ago [ - ]

Tbf, most of the "real benchmarks" have issues that are just as bad. Assessing LLM performance is just hard

oceansky 7 hours ago [ - ]

And personal too. Different engineers are using them for different use cases.

ramraj07 3 hours ago [ - ]

The important point is that your benchmark is pretty much irrelevant for the actual usage. Thus whatever conclusion you draw is not just irrelevant but misleading.

meander_water 12 hours ago [ - ]

Thanks, I didn't mean to be brusque, but I have seen a lot of these vibe tests lately that come to grand conclusions like "X model is better than Y" from the result of a single prompt.

Appreciate you sharing the results of your tests though!

jameswhitford 11 hours ago [ - ]

I appreciate the feedback!

esperent 12 hours ago [ - ]

On the other hand, I did just leave my pi agent running GPT 5.5 overnight on a clearly defined, long running task. It's been running about 10 hours now and it's mostly done. So this kind of use case is also valid.

Thinking about it, I would say that the majority of agentic work I do, by a long shot, is subagents which are launched from the main session, using a prompt of its choosing. Those could be considered short versions of these fully autonomous tasks.

thunspa 8 hours ago [ - ]

Care to share more about your pi setup? I've recently started using it (after long-time Claude Code work) and was wondering how you'd achieve these long-running tasks. Do you allow it to spawn sub-agents? Thank you!

esperent 7 hours ago [ - ]

My pi usage over the past ~5 months went roughly like this:

* Install pi and a bunch of extensions from their package repo

* Realize that all the packages (with a few exceptions) are massively overcomplicated and vibe coded

* Ask pi to rebuild a very simple version of the packages I used. So e.g. subagents - all the default subagent extensions are massively complicated with named agents, recursion, communication. I made one that stripped all that out.

* Then whenever I hit an annoyance, spin up a parallel session and fix it.

It's less work than it appears because I have ~5 extensions: hooks, subagents, background processes, a custom footer, a loop command... Maybe that's it. Within a couple of days you can have a setup pretty close to Claude Code but with a fraction of the base context use. After gradual improvements over a few weeks/months you'll have a system far better, tuned to your exact preference.

Of course, just like Linux or any other highly tunable system equally important is having the restraint to not spend all your time tuning it. I've definitely had a couple of days where I was bored with my real work and did that, but whatever, it beats browsing reddit.

As for getting long running tasks, I set a looping message every ~20m and tell the agent to strictly track progress in a session doc, then reread and continue after each compaction.

ethanpil 44 minutes ago [ - ]

I'd like to study your setup. Would you be willing to share? Perhaps a github repo of your 5 extensions or even a pastebin if you would be so inclined. I would be grateful to learn more about this by studying from your success...

ijidak 5 hours ago [ - ]

What type of task are you running for ten hours? Is this a programming task?

I've not come across a programming task that would take an LLM ten hours.

nfriedly 2 hours ago [ - ]

I'm not the person you asked, but if they're running in their own local hardware, then it might just be a lot slower than what the big providers run their models on. System RAM is a lot cheaper than VRAM, especially if you bought it last year.

jameswhitford 11 hours ago [ - ]

Yes, part of the reason I chose the one-shot test was really to test long-running tasks. A lot of people seem to be experimenting with this format, for example in the now trending loop-writing workflows. And really I am interested in diving into the murky waters of these novel workflows.

segmondy 2 hours ago [ - ]

One shot prompt means you give the model and input, you get an output done. This was not a one shot prompt, but an agentic task as shown by the tool calls.

ritzaco 13 hours ago [ - ]

sure that's why we look at a mix of formal benchmarks, one longer analysis of a side-by-side, and various other people who we trust to form an opinion, all covered in the article - not intended to be a formal benchmark, there are enough of those.

patates 12 hours ago [ - ]

Then maybe you should add that caveat emptor to the article?

You make a very strong claim at the end that the hype is mostly real, and making it clear to what extent your claim holds should help the reader.

12 hours ago [ - ]

[deleted]

unliftedq 12 hours ago [ - ]

Totally agree, a single one-shot prompt can't prove anything.