Honestly, I don't feel the least bit of excitement here and I'm normally enthusiastic about AI.

Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.

I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.

You’re correct, it’s a great solution for a very narrow set of high volume classification needs that require very low latency. But that’s it.

I keep evaluating Jev for my product because of the hype, but the reality is that for my tasks, Luna is more accurate, only a little more expensive, and the latency doesn’t matter. I’d rather spend the extra $50 / month or whatever than have to shoehorn in another API and provider, and also lose the ability to change reasoning level and get reasoning summaries for eval purposes.