Also open sourced a bunch of infra to go with it.
Anyone who claims open source and open weights models are "decel" needs to get their head checked
https://github.com/MoonshotAI/MoonEP
Also open sourced a bunch of infra to go with it.
Anyone who claims open source and open weights models are "decel" needs to get their head checked
https://github.com/MoonshotAI/MoonEP
This comment would be much better without the second line
I'm not sure I understand the case for open-source models being decelerationist, is this it?
Decel:
- Potentially reduces investor appetite for funding big labs.
- More risk of powerful AI getting in bad hands -> more regulation.
Accel:
- More competition so big labs can't rest on laurels.
- More research in open, so all labs can accrete advancements faster.
I feel like open-source = acceleration has a much more clear argument. (and how bad would deceleration be in any case?)
I think the argument is that decentralization leads to deceleration because it means less centralized funding and data. Those are the two primary ingredients for accel.
The problem with the decel/accel rhetoric is that it lacks nuance.
I think it's basically open weights => more inference competition => less profit from inference => less training competition
> less training competition
I think you meant less research and experiments in big labs because they don't get all the AI money.
Training is expensive, but they also have more than 10 000 of employees combined and they cost a lot of money.
If your worldview is “most of the progress is made by closed labs, then open labs fast-follow” (which isn’t implausible given the documented distillation of Fable), and further that open labs cannot make make meaningful progress vs the closed labs except by fast-following and that they won’t pick up the ability to make progress after the closed labs are gone, then driving closed labs out of business slows down overall progress.
I think it's pretty hard to hold that worldview: Anthropic couldn't ship a reasoning model until they copied DeepSeek R1's homework, and they've all copied DS-style super-sparse MoEs at this point too.
With slightly different cherry-picking, you could equally well claim that DeepSeek couldn't ship a reasoning model until they copied the idea from OpenAI's o1-preview, and they also copied MoEs from Google Brain/Jagellonian University https://arxiv.org/abs/1701.06538 way back in 2017, too!
But ultimately these were ideas floating around in the air, if one group hadn't done the experiment, someone else would have.
That’s a really good point. Folks really need to read the papers coming out of these Chinese labs. Every paper from the DeepSeek team has been a step change.
Open Source models decelerate growth of closed AI. For people who think (or want) AI = closed_AI then that argument has weight. Good luck getting them to update their priors.
the argument is that we should all fold and let sam altman burn trillions of dollars on naive scaling and pay monopoly prices for their closed APIs until the models are good enough to be closed off for "safety" reasons so that they can take an even larger cut by competing directly with us
Open source AI is actually a lot less "powerful" than genuine frontier models, i.e. it has a much tighter inherent capability ceiling. This is "decelerationist" from a purely AGI-pilled point of view but it's actually great if you're worried about a capabilities arms race putting AI Safety at severe risk.
Kimi K3 is plausibly a lot less dangerous than a totally jailbroken ChatGPT/Gemini/Claude Sonnet (let alone Opus or Fable!) and it's quite deeply weird how no one seems to be calling for those models to be banned or restrained by further regulation. Why the double standard against the less concerning (but more efficient!) open weight models?
Do you think they are inherently less powerful? I'd imagined that closed labs have a head start / more funding so the open labs are playing catch-up.
Is there a world where open source models end up at the frontier, or do you think there are structural/first-principles reasons why this won't happen?
If you're targeting widespread local/on prem deployment which is what many open weight models are doing, that inherently limits your scale in terms of total model weights/inference-time compute compared to running in a few centralized datacenters. A centralized model will always be able to leverage a larger scale of deployment, placing it much closer to the genuine "frontier".
Why? Absolutely correct, especially considering the position of the person they're referring to
It’s a different topic but it’s correct imo. Open sourcing things is the best way to accelerate development.
Put another way, if you want to slow things down, put it behind a paywall, tag ideas ans “intellectual property” (meaning you’re the only one who can use it) and get the lawyers involved (injecting our slow legal system).
None of the above is a judgement call on whether development should be accelerated.
Why?
[flagged]
I know it's a hard ask on this site, but I need you to start parsing content and not tone. It was a helpful bit of context, even if it was a bit vitriolic.
The content was 'you need to get your head checked'. That isn't tone, that directly implying that if you hold that position, there is something wrong with you. It's rude and unnecessary.
Why should I waste time parsing content and not tone? Why can't the commenter just avoid the tone? It even saves time since you can write less!
Because discourse around tone isn't productive. You could wipe this whole comment thread, starting with the parent of mine, and lose exactly zero information.
[flagged]
> because it gets people like you going
So you're admiting you were trolling?
no, the goal was to spark a conversation about the value of *open AI* and it looks like it worked
Some battles are simply not going to be won.
I, for example, dislike reading comments complaining the submission (or another comment) is LLM generated. Focus on the content, not the style.
I'm not going to have my way, and nor shall you.
Oh, I understand. It's the nature of the voting system of comment feedback. Reddit behaves the same way. Arguments become competitions to see who can inject enough vitriol while still maintaining a placid genteel demeanor. The one who "wins" (convinces the peanut gallery to upvote them) is the one who can avoid looking like they got mad.
It's why people can advocate for ethnic cleansing here, and that's fine as long as they word it correctly, but if someone calls them an asshole about it, they're flagged.
Really common in rationalist circles from my experience as well because they believe that true statements aren't always normative, and that, since their arguments are true because they're rational, their statements aren't necessarily normative. Begging the question, of course, but I see that in situations like Scott Alexander's defense of "human biodiversity" theories (the whole HBD moniker is itself an example of everything I'm talking about condensed into two words).
> Oh, I understand. It's the nature of the voting system of comment feedback. Reddit behaves the same way. Arguments become competitions to see who can inject enough vitriol while still maintaining a placid genteel demeanor. The one who "wins" (convinces the peanut gallery to upvote them) is the one who can avoid looking like they got mad.
Depends on which subreddit and which flavor of groupthink. The behavior you say is upvoted is one I often see downvoted to oblivion on Reddit.
> but if someone calls them an asshole about it, they're flagged.
That's because name calling is against HN guidelines.
Just an aside since it's not clear: I think the asshole is the one using labels like "decel" (or even "MAGA" unless the person self describes). It's irrelevant if I agree with the rest of the comment.
It's OK to call people out for name calling while still agreeing with the rest. It's problematic to require one acknowledge the quality of the rest of the comment when calling them out on their name calling.
Put in a less twisted manner: If it's OK to address his comment sans the "decel", it should also be OK to address his use of "decel" without discussing the rest of the comment.
It depends on if we expect posters to be informative and include nuance.
To me its clear that it is decel
the only reason other labs can catch up is because the frontier labs can be distilled, and they siphon a % of the labs' revenue to reinvest into the next iteration
full accel would mean nationalizing the big 2 labs and locking in manhattan project style until RSI
(Edit: some great counterpoints in the replies. my view has definitely been changed!)
only if you only get your news from main stream business press and Big Lab propaganda channels
There's no chance K3 is a distill of Fable, it came out way too soon after the limited fable release to be feasbile.
If you look at all of the top ML conferences, chinese labs contribute way more to advances in ML than "Open"AI and Anthropic: https://www.reddit.com/r/TheMachineGod/comments/1pi4q7f/pape...
This K3 release just helped every other lab on the planet stay in the race by making it possible for them to build on top of it, placing them at the frontier starting line instead of having to spend billions of their own dollars and risking it all to attempt to catch up.
The open source contributions I linked to above will move the whole field forward and reduce the costs of training and inference for everyone.
Open science compounds on it self, every new advancement pushes the field forwards and opens up new grounds for future improvements.
It is impossible for a single closed lab to consistently stay ahead of the rest of the field, especially in a huge growing research area like machine learning. The only only advantage the big labs have is money, but the naive scaling game is not sustainable long term when you have to pay 10-100x more then the fast followers and we start getting more and more open models or use case specific models that can handle 90% of high volume use cases.
Research is a high variance, low expected value activity, meaning that the few large concentrated labs have to be conservative with their bets and double down on proven things when scaling up. The rest of the field is like a diversified portfolio, with thousands of players making smaller riskier bets that only require a few of them to succeed (like K3 did here, and DeepSeek a year ago)
EDIT: also if you look at most of the work from OpenAI, it's mostly taking existing promising open research work and scaling it up. (except for things like CLIP and etc from Alec Radford)
Appreciate this point. Back in grad school, I published a peer reviewed paper with all the source code and datasets used. Got heckled at a conference talk by staff from a commercial lab. They shouted, we figured this out five years ago, lol. Also our approach is still better. But they don’t release their source code or publish much, so no one knows what this approach is or if it’s actually better.
And years down the line, lots of other research labs used my code and cited my paper.
Oh this definitely happens all the time. I was an early employee at Clarifai, which won imagenet a year after alexnet and we were able to stay on the frontier for about 2 years before a bunch of open source models were matching our results. It was always some random PhD research project spinout or some random kid in Boston named Alec Radford.
We had a bunch of things that we never published that ended up being major research findings years later at top conferences.
OpenAI's head of strategic futures publicly stated that you can't explain the quality of the newest Kimi via distillation.
Further, you can just read the papers released alongside most open models. Plenty of hugely influential research results published that drive the frontier forward. It's not like these models are just existing architectures downloaded from Huggingface and trained on frontier lab APIs.
Distilling can be done by fast follower closed models too, so this argument against open models doesn't hold up.
being able to replicate it in the open means that there's nothing special about frontier models.
Frontier models would have to do something extraordinary or unique, or unreplicatable, because clearly there is no moat, and US companies are sitting on huge nvidia valuations and get surprised when competitors beat them.