Is there any LLM from exactly one year ago that would be worth running?

In Aug 2025 you had

- OpenAI o3

- Opus 4.1

- Gemini 2.5 Pro

- Grok 4

Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free.

Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.

> Is there any LLM from exactly one year ago that would be worth running?

Bad perspective: consider the correction: "when are thresholds of sought quality reached"? Hence: not "is there a 10yo from last year that could compete with the current 13yo", but "will there be a 30(?)yo from last year that could compete with the current 33(?)yo" ('(?)': the scale of yearly growth in the future is uncertain).

It's not just about it "being smart enough". It's about there being actual user demand when it needs to compete with the shiny new model.

A 10 year old iPhone is probably good enough, but is there demand for it? In a vacuum a 10 year old iPhone is good, but why would you pick it if you can have a current one for a reasonable price?

Gemini 2.5 Pro was very good at writing single, somewhat complex functions. Sure, the rest of the loop would still take time, but nearly-instant implementation? Sign me up.

We still run GPT 4.1 for some of our use cases. We want to replace it but are having trouble finding models that are as fast with similar or better intelligence.

There's nothing fast about GPT 4.1. It's ~50 tps AFAIK. Of course it doesn't use reasoning, but you can run modern models without thinking as well. GPT 5.6 Sol without reasoning should destroy it in intelligence.

That is fkin wild. o3 was just a year ago? The progress is truly insane.

Yeah I had to double check, o3 feels like it was ages ago. But GPT 5 came out Aug 7, so it's only one day off from my 1 year ago cutoff!