> Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.
> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.
Early access, no weights no tech details, just a sign up here for info
I'm all for more open models, but talk is cheap and this is a rather pointless announcement without anything backing it up. Publish your weights and HF repo or shut up IMO.
Is that all that a company about to give away the product of 10,000 GPUs running for a month gets to be now? give it away without a single promotional post or shut up? I support open source as much as the person but this is pretty caustic.
I am also starting to take issue with "we've developed a new cutting edge model, and nobody can use it" announcements. Most recently with Google's Argon, at least they were using it internally and they'd be slowly rolling it out. This one is from a company I've never heard of, they're not releasing weights or offering API access at this point, it feels like a fairly worthless announcement.
It says in the very beginning that everything will be opened up by the end of the month. Maybe slightly annoying to you but not worthless. I for one am very excited to see some new American open models since we have called so far behind the Chinese. Looking forward to trying it out.
Yes -- because they're late to the party (high performing open weights models have been a thing for a couple of years now), and therefore will be compared with all the other open-weights model providers they are competing against.
It's not just that they are doing users a favor with weights; they are just as much seeking favors with attention and usage (in a crowded market!).
There have been previous "we will release the model" when the model release never comes in the past.
I think the comment you are replying to is unnecessarily hostile too though.
> give it away without a single promotional post or shut up?
I think the point is that people are happy to see promotional posts when they actually release it, but only then and not before.
Unfortunately pinky promises from corporations to release something at some indeterminate time in the future aren't worth the bytes they're stored in, especially in the AI industry which is full of grifters and charlatans.
The open weight community really is an odd one. Millions of Dollars for pre and post training given away for free and most often with very permissive licenses that allow commercial use (be it US, EU or mostly Chinese)
Yet going by the comments on localLlaMA or HN, those companies are the devil :-D. Colour me surprised.
You know, the other night I trained a model on my secret stash of GPUs, that now outperforms Opus 5.5 on pelican benchmark and Jev on classification speed, while running on a potato.
Will release weights soon.
I don’t understand what is wrong with some people lol. It says in the very beginning everything will be opened up by the end of the month. It’s like a bunch of toddlers throwing their milk cup on the ground and pouting because they can’t have their new toy right now!
In the case of Facebook, no good thing they do will ever undo the evil things they've done, let alone the evil things they're still doing right now. How convenient for the devil that a single act of "charity" should make him immune from all criticism. By all means, praise whatever good Facebook does in the world, but don't kid yourself about what Facebook is.
You don't get claim to be open and then not release your actual product.
Show me the mone^H^H^HWEIGHTS
They claim that they will release the model as open weights later this month.
That means that they have the 31th of October as the deadline to make true their claims.
The fact that they give early access to some may mean that they want some beta testers before the public release.
And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none.
Not sharing the data is pretty standard because 1) it tends to get the lawyers involved and 2) good data is critical for getting good results.
Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.
This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.
This isn't true.
Companies pay lots of money for proprietary agentic trajectories which are used during RL. These are things like "Task: summarize stock levels for months end accounting" which then traces the task though using SAP to look at different SKU stock levels, exporting them and generating summary Excel spreadsheets.
This is very different to the "scrape the internet" datasets that a table stakes for training a LLM.
Xiaomi released a fairly developer-centric dataset like this here: https://huggingface.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
SpreadsheetRL is another fairly specialized dataset: https://spreadsheet-rl.github.io/
Data has copyright issues, so one can't share it generally without getting permissions from all of the copyright holders. The data is not theirs to share, anyways. The derived (learned) weights are a different matter.
True for the pre-training data scraped from diverse sources. Less so for the later stage data for RL which is by all account more of a differentiator. In most cases the labs themselves produced the data so they are the copyright holders (or they are borrowing it from other labs via distillation). A lot of it is synthetic data, and since you can't copyright AI output, it becomes less about copyright and more about trade secrets.
If I train a model on a song's lyrics, and the user asks the model to recite the song, and it does, so in RLHF I downvote the generation because it refused-to-refuse to regurgitate the song's lyrics... is that RLHF session copyright-encumbered?
this is very normal for frontier lab companies. you need good data either synthetic or labelled (all the chinese open source models have their own armies of data labelers)
> Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model.
> We will release the weights, technical report, model card, and developer artifacts later this month.