Decision models have the potential to have an even larger impact on the Real World than LLMs have to this point (which is obviously quite large). But the model itself matters less than the product experiences you build around the model, and its very likely that the incumbent labs are treating the area as something more like "oh yeah I guess we can ship that and then forget about it" rather than investing in what building business processes on decision models looks like. Unlike full language models, I don't think the primary business of Typesafe will be serving Jev at API pricing; it'll look a lot more like putting Jev at the center of a much more expensive suite of software.

There's the potential for an inverse LLM play. In contrast with LLMs, all that seems to matter is the model, and the products the labs build around the models are all really samey and boring; the same left panel list of agents, main view agent conversation, right hand extra context, and we're now in the era of everyone creating the same cutesey furry friend on top of all this tech.

> But the model itself matters less than the product experiences you build around the model

This is a very important insight. And it applies to LLMs as well. Very few people were impressed with the capabilities of GPT 3, it was mostly a techie novelty

But then when they added chat on top of gpt 3.5, all of a sudden it was a huge hit. Sure there were improvements in the model from 3 to 3.5, but the biggest impact was from the chat experience

Conversely, when they created Eliza, a basic chatbot more than 50 years ago, people even got addicted to it, despite it’s ai model being something super rudimentary and basic compared to what we have now. The model capabilities didn’t matter as much as the experience the chat created

> Decision models have the potential to have an even larger impact on the Real World than LLMs have to this point

Why?

If youre the one or have spoken to someone implementing "AI solutions" inside large companies recently, a decent chunk of it is soft policy enforcement with very basic context. They moved from gemini 2.5 flash lite type models to jev. Which is also why I found the price comparisons to "GPT Astra" on Twitter rather funny.

Yes, the best way to think of a general classifier like this is like a smart switch statement. Essentially a "JEV" like thing becomes a sort of programming primitive. Once you see it, it's hard to not get excited.

But even if others surpass them and make better solutions, the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.

I sound like a fanboy but I swear I 'm unaffiliated with typesafe. I was building my own version of this way before they announced JEV (mine was ALE and it was mentioned here on HN for a bit), in use for VR gaming (so one can give commands to NPCs with voice and supports multiple commands in sequence in a single pass), but I missed the "killer usecase" of being a new primitive, like everyone else.

TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again

> [...]the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.

Counterpoint: there are a lot of one-hit wonders, and they vastly outnumber the idea-factory people. This is not to minimize those people, a single idea can be very successful (see Zuckerberg), but it doesn't mean your subsequent ideas will also be great (see Zuckerberg)

I remember your blog post! Thanks for writing it, was pretty cool and a practical application.

Here if anyone is interested: https://pantel.is/projects/ai-gaming-companion/

Honestly, I don't feel the least bit of excitement here and I'm normally enthusiastic about AI.

Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.

I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.

You’re correct, it’s a great solution for a very narrow set of high volume classification needs that require very low latency. But that’s it.

I keep evaluating Jev for my product because of the hype, but the reality is that for my tasks, Luna is more accurate, only a little more expensive, and the latency doesn’t matter. I’d rather spend the extra $50 / month or whatever than have to shoehorn in another API and provider, and also lose the ability to change reasoning level and get reasoning summaries for eval purposes.

>> TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again

This all sounds intelligent and likely, and yet we can come up with countless counter examples where the first mover is not the big winner, and nobody cares about who did it first. There is typically way more value in nailing the execution of a big idea someone else came up with, rather than "seeing the future".

> Unlike full language models, I don't think the primary business of Typesafe will be serving Jev at API pricing; it'll look a lot more like putting Jev at the center of a much more expensive suite of software.

Precisely this. Should be top comment.

Also, the model moat is understated as training data for these purposes also accrues to the winner, which due to the first mover advantage as well as the distribution advantage you speak of, is typesafe. In contrast to relatively open coding data. Openai anthropic also have that, but like you say its a different business.

The are still just 1tok output of pretty standard llms just along with the logprobs converted to some json

At least in theory (TypeSafe has been pretty close-lipped about the details so this might just be hot air, and I think the evidence is a bit spotty) this is false, since they use a different reinforcement training method.

If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.

[dead]

Yeah, agreed. A model or primitive on its own has no moat and frankly limited value. The paradigm behind "System One" models on the other hand is potentially huge.

https://seldon-ai.com/blog/fronter-llms-are-semantic-interpr...

Skimming the post, it seems to argue for reconstructing the very rigidity that LLMs let us escape, and that very aspect of LLMs is what made them useful and explode in popularity so much.