Honestly, I don't feel the least bit of excitement here and I'm normally enthusiastic about AI.
Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.
I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.
You’re correct, it’s a great solution for a very narrow set of high volume classification needs that require very low latency. But that’s it.
I keep evaluating Jev for my product because of the hype, but the reality is that for my tasks, Luna is more accurate, only a little more expensive, and the latency doesn’t matter. I’d rather spend the extra $50 / month or whatever than have to shoehorn in another API and provider, and also lose the ability to change reasoning level and get reasoning summaries for eval purposes.