Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).

Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).

Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).

Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.

Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).

Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.

Well, Jev doesn't meet any real compliance requirements but OpenAI's models do. So even if Jev is faster, any customers with compliance needs will obviously pick OpenAI because they can't pick Jev out of necessity.

Or even just- they've already procured OpenAI.

That's a big motivator.

They can follow up in N weeks with a better one. Even if your eval is true and it’s worse, planting a flag makes sense. Some people will just use OAI because it’s OAI. No one will remember their week 2 evals in a few months.

I think releasing something like this makes sense even if it's underbaked, it's still very cheap, and if you have existing enterprise OpenAI relationship it's a lot easier to onboard something like this than set up a new Jev contract.

How these models play out is an open question but existing provider contracts and T&C are important for enterprise.

Good point, commercially, being an existing partner is always easier for adoption. Heck, why I'd like to get Mercury Decide to replace Jev myself, rather than one than two to work with.

Still surprised they even leveraged Luna for this. Given their resources in data, compute and manpower, would training a decision model from scratch take that much longer to not make sense given the cost, compute and performance advantages that would likely provide?

3 times more expensive at twice the latency with lower performance is a tough sell, though yeah, prior relationships will likely smooth some of those deficiencies over.

yeah and this isn't a long term solution, stand this up, see what value you get out of it, and in a few months you cans witch to whatever the best decision model is

[flagged]

i would be shocked if luna decides is less generally capable than jev. jev has failed to understand any novel domain I've given it. i have found that i use it only when "some data is better than no data"

Very task-dependent of course and mine are unique to say the least, so could see Luna being better in certain domains, even if my measurements have not shown that yet, happy for anyone to show otherwise.

For what it's worth, ran every task twice on each model, most were for some UI component synthesis and charting insanity that is a bit hard to explain, but some were simple tag selection, basic noul at threshold 60%. Essentially, whether to use the provided tag given the title of a browser tile:

  1. Title: "Mortgage calculator: estimate your monthly payment (Bankrate)"
     Tag:   "house hunting"
     Jev:   0.69 (yes), 0.72 (yes)
     Luna:  0.21 (no),  0.21 (no)

  2. Title: "S&P 500 index: live chart and news (Bloomberg)"
     Tag:   "investing"
     Jev:   0.91 (yes), 0.90 (yes)
     Luna:  0.56 (no),  0.56 (no)
Of course, tags can be a bit subjective, but in these cases, I'd argue the values provided by Jev were far more representative of my subjective assessment over Lunas. If SnP stuff on Bloomberg isn't investing, nothing is.

Goal for tagging is mainly a near instant, over writable, sane default provided to users in the background. Resolve the whole "I love using Notion/Obsidian/PKM software of your choice but spend 80% of my time just thinking about the ideal tag before starting to read" issue. Lunas output is not really helpful here.