Disclosure: I work at H2O.ai.
We released an Apache-2.0, open-weight 4B decision model that scores above Jev 1.13 on JevBench's composite score (72.5 vs 71.5) and is currently the top open model there: https://benchmarkheaven.com/jev-models . Newer models coming even larger than beat Jev in intelligence as well.
- Same contract as Jev: state + typed questions in, calibrated probabilities out, one forward pass, no generated tokens. - Your data never leaves your environment, and there's no per-call fee.
Weights, card and run instructions: https://huggingface.co/h2oai/h2o-lightning-4b
Is it possible to quantize the model? Don't think I want to run 8 gigs on my local just for decisions but curious about it for sure
Update: Microsoft just announced Microsoft-Decision-1 (a post-trained Qwen3.5-9B) and benchmarked it against H2O-Lightning-4B, among others: https://commandline.microsoft.com/microsoft-decision-1-model...
One note on their speed comparison: the JevBench board shows self-hosted models with an adjusted latency of "2x + 0.15 s (assumption, not measured)". Our measured p50 is 29 ms; the adjusted figure is 0.21 s. They report 85 ms for theirs.