When you say "trained classifiers", you are referring to models which are trained (or fine tuned) to work on one specific problem, right? That is the opposite of "general purpose".

Would Jev be more accurate in a specific task if it had been developed only for that task, as opposed to general purpose? Of course it would. So, sure, Jev is trading accuracy for generality. According to you "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier.

> work on one specific problem, right? That is the opposite of "general purpose".

A business doesn't need Jev for the sake of Jev. Most business are solving specific problems.

And fine-tuning got a lot cheaper these days - I've seen claims here on HN that ~500 examples is enough to beat Jev.

> "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier

There wasn't. The hope is that there was a latent demand, but we've yet to see if it's truly latent or just manufactured.

Noone is saying "hell yeah, finally we got a general purpose classifier, my business needed it so much". The typical message is "this seems cool, let me see where I can apply it".

The fact that name itself is a play on Jevons Paradox illustrates that there was no demand until Jev was released.

Yes, businesses are solving specific problems, but most businesses have more than 1 problem to solve. No, it is not economical to pay a data scientist to develop a custom model for each of your tiny problems. It is often much more economical to use a general purpose solution, like an LLM, or now, Jev.

At this point there's no need to pay a data scientist. You can literally ask Claude to do everything for you: extract real examples, classify them, post-train a model, and ship an API.

Now, it could be viable if your business has literally hundreds of problems thats require classification. I just haven't seen those.

I treat the fact that almost noone was doing that as evidence that decision models aren't that useful/groundbreaking. That, and the fact that every single demo I saw was either fake (e.g. playing games), contrived, or plain wrong (e.g. using Jev for compaction).

No, we're not at the point where you could ask Claude to do all of that, unless you have a super easy problem to begin with and/or you don't care about output quality. Feel free to link a counter example.