I'm interested to see where they want to land performance-wise (i.e. which point they choose on the scaling curve) and the niche they want to carve. They have a decent ways to scale beyond trinity large, in paticular on posttrain/RL before they are competitive with open-weights, especially internationally.

Deepseek is explicitly banned [1] at LLNL and I wouldn't be suprised if there's a blanket ban on all Chinese models. But nowadays models like tera/luna could fill this area of the pareto front, and LANL already runs openai models on their clusters [2]. Maybe it's in custom SFT/RL, for instrument control or sensitive topics? But you'll still have to compete with frontier models + a harness.

I would have also liked to see a carrot tied to their offer. It'll be hard to get teams to contribute RL gyms or curated text. But throw in a "we'll fund a postdoc/student to do that" and I think you'd have teams scrambling to apply.

[1] https://hpc.llnl.gov/about-livermore-computing/ai-ml-lc/lc-l...

[2] https://www.energy.gov/nnsa/articles/nnsas-los-alamos-nation...

That's interesting a locally hosted LLM would be banned. I'm assuming locally hosted is included. Do they think it's been trained to sabotage equipment?

I'd actually suggest a great starting point would be a local command reviewer LLM. Could ostensibly be a modern AV type thing. Particularly seeing this lately has driven the need home deeper to me: https://x.com/chrisbanes/status/2085341561609425230?s=20

An open weight tool call auto-reviewer, has all sorts of achievable scaling curve milestones.