This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric.

Am I missing something?

> Am I missing something?

It's pretty clear from their framing ("Beam advances the Western open-weight frontier") that one of their main selling points is not being a Chinese lab.

I can't imagine that mattering to many individuals, but I guess someone out there has a government contract that forbids the use of foreign models

Reflection raised on the idea of creating the "American Deepseek Project"

Who's funding this?

Looks like Nvidia, Eric Schmidt, Sequoia, and a host of others https://techcrunch.com/2025/10/09/reflection-raises-2b-to-be...

I still don't get why, after over a decade on HN, people refuse to Google very simple questions.

Reciprocally, people probably do go find out. The great filter is also how many of them come back to post the useful interesting information.

I still don't get why, after 25 years of slashdot, still ask why people DRTFA and complain about it as meta commentary.

InB4: kids these days :shakes-fist-at-cloud:

Sometimes we just enjiy having a conversation with people. It's how we got many answers before Google exists. I believe such choices have many, positive effects on people that society is losing.

quesiton is more like "lets analyze the motivations behind funding this"

And a much better post would be "Just checked, and X, Y, and Z are funding this. My bet is that Y is funding it because $REASON..."

Asking an easily-searchable question is just lazy.

Profit. Always.

because my templeos goes straight from ring0 after bios straight into a ui for hackernews that only lets me scroll, click into comments and type comments.

Try using the same Internet that you used to download templeos to search for stuff.

I don't get why anyone posts anything on HN when they can just have an LLM generate an entirely self-contained conversation now.

/s

The conversation is why.

Multiple independent approaches are cool and all but fully open source model training (datasets, pipeline, checkpoints) should be taking advantage of being open and share runs/budget between different entities.

We're still at the stage where every new entrant is welcome in my opinion. Doesn't need to be record-breaking upon initial release.

It depends! If a startup is entering with a large model to face other larger models, it must be better at least in 1 meaningful dimension.

500B params performing worse than other OSS of the same size is pretty meaningless if no one will use it.

Look at meta, while they went from open to closed, they got better as time went on

I mean, the obvious way it's better than the Chinese open-weight models is literally the fact that it's not Chinese and is therefore less likely to be banned or restricted. Chinese models cannot be used on certain government systems already in the US, and regulators are actively considering expanding these restrictions more generally (such as adding it to the Entity List).

I disagree. Sure let them play and see if they can improve. But this model has more compute and more training data than the predecessors it fails to surpass. That only means their training regime is inferior if their predecessors did so much more with so much less. That inferiority should not be encouraged.

You don't just magically do better than everyone else on every metric on your first go at something. Doing worse than others and refining is how pretty much everything works.

I think i captured that in my first (second?) sentence

The reality is they trained a model and it looks worse on benchmarks than Qwen or GLM. I don’t see how sharing the weights hurts anyone? Even when Llama 4 came out and it was a dumpster fire, it didn’t affect me personally.

> That only means their training regime is inferior if their predecessors did so much more with so much less

Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.

None of that means they shouldn’t release their model.

Them releasing the weights doesn't hurt anyone. It's the peanut gallery clamoring to put them onto the same pedestal as actual tier 1 companies simply because they aren't named openai or anthropic that is hurtful.

I wonder if that's an indication that they are not distilling which limits how good they can get.

Openai, grok, and Anthropic aren't distilling. Theyre just second class. It's not a big deal, we just shouldn't be lauding them for being second class.

> Openai, grok, and Anthropic aren't distilling

Says who? We know Grok does at the least. They admitted it openly.

Musk said under oath that they use distillation for Grok.

And it still sucks. My apologies to the Cursor team but that's just very very poor performance.

Alternative explanation is that the Chinese have far more technical talent than anyone else, along with the infra and capital to build out these models.

My money is on the latter explanation, tbh.

Reflection is explicitly marketed as the 'US' DeepSeek

seems like they are aiming to provide both inference and RLaaS for american companies and western govts. even if they never fully beat deepseek if they get close enough the fact that they're American will help them close deals

Apparently there is more to making good models than copying everything on the internet.

deepseek is from the evil east, this is from the virtuous west

12 yards long, 2 lanes wide, 65 tons of American Pride! Canyonero! Canyonero!

Yes it's (hopefully) not distilled from every single major American provider.

new entrant in this weight class, US lab.

Beam goes brrrr