The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.
With open source projects, the benefit was that each individual could improve the complex system (e.g. Linux Kernel) interpedently, and over time the benefits accumulated. With models right now, there is just no way to do distributed training, or really, any large scale parallel way to improve them.
So whatever the short term strategy driving publicizing the model weights (e.g. potentially, to create a price war in order to put pressure on western companies and deprive them of the money they need), we can't ignore the fact that incentives and decisions could easily change in the future, and unless there is a way to truly decentralize models improvements - the party could stop at any time.
Complements.
But different huge companies have different incentives. It is very much in Nvidia’s interest to have me running a powerful open source model on a $4k machine that they sell me.
Is it? When they could be having you running an even more powerful model on a $50k machine they sell by the pallet-load to enterprise consumers? We already see RAM manufacturers abandoning the low-end market in favor of server support. It's not clear to me that Nvidia sees personal GPUs as their best long term investment compared to selling millions of server-farm class machines
You mean a $500k machine, or a $15M rack... the costs have gotten unimaginably large from the lens of just a decade ago.
... for AI... which people have to convince themselves they 'need'. I am a user, I have a 1m$ rack for my small company for AI, but I have currently 189 racks that cost about 15k one off per rack. And those make us a LOT more money than the AI ones ; most are departmental apps or websites that take very little processing power and are low on everything while people pay nice money to keep them going. Our support is better than the rest; you can just call me or come to my house and we'll do something. Not many clients leave, ever. For the past 30 years. I buy 5-10 year old servers which are complete overkill for 99% of my clients (some are fortune 500 companies but departmental stuff) and they just never leave. Makes vastly more money than AI and we are still cheaper than AWS and the department does not have to get a flexible budget (like they would have to if AWS) with us. We just have the same fixed price + yearly inflation correction and that's it. For 30 years.
Selling to individuals can be a hugely more robust predictable business, the problem with selling by the pallet load is spiky revenue that can also quickly fall off a cliff if larger customers stop buying. The other problem is sales negotiations driving down margins for bulk buyers etc. Consumer hardware is a very attractive market in lots of ways, just look at Apple.
Right. It's probably important to distinguish what the hardware manufacturers' incentives are when their supply is constrained and when it isn't. (I am not an expert on anything.) As long as supply is strongly constrained they're naturally going to sell to the highest bidders, which are the big LLM SaaS players. If and when supply is no longer constrained, though, things will look very different:
1) The LLM SaaS companies are a form of vertical disintegration for the hardware providers, a middleman covering costs and taking profits out of the money that comes from customers to the hardware providers. That changes somewhat if there are no longer good models available for local use at no cost to the hardware guys, but only somewhat
2) The LLM SaaS companies are efficient users of their hardware resources. While supply is constrained this helps to make them top bidders and so attractive customers for the hardware manufacturers. When supply is not constrained this should reverse. Which is the more attractive class of customer to a hardware maker: the company full of people with higher degrees who spend their whole working day fighting to pare back resource usage, or the guy who leaves his laptop idle about 18 hours per day on average?
It's notable that nVidia, for instance, has continued to put significant emphasis on AI compact desktops and laptops. And while no doubt that's partly in the service of better developer relations and good PR in general, it's probably also nVidia eyeing the exit, and preparing for a future transition from selling shovels to the army to selling shovels at Walmart. But of course the future isn't clear and obvious. If the hardware makers, maybe the RAM guys in particular, turn out to have underbuilt future capacity starting in the present then we could be stuck in constrained supply for quite a long time. (Futher) government action could affect things etc. etc. And if the frontier labs soon find new ways to use still larger amounts of memory, GPU capacity etc. that isn't butting up against diminishing returns then they'll likely remain kings for some time, though that does not seem probable now.
They might not release such a thing, lest it puts all of their customers out of business.
Who funds the majority of cutting edge scientific research?
Is it companies or is it governments?
If governments around the world see LLMs built from public knowledge as pre-competitive as the public knowledge itself, then why wouldn't they sustainbly fund it?
There's a lot of individual effort of improving the models. See how many finetuned models and LoRAs are there on Hugging Face.
fine-tuning a model is very different from training the whole model in terms of resource requirements.
Soo? If they need a better version, the world can pool resources together, form a company that trains the model, then the company goes under and the model becomes open source again.
> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.
Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."
There is more to society than capitalism.
> Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."
> There is more to society than capitalism.
I don't read GP like that. I read it as "we should recognize a situation of unstable incentives for an important outcome, and start thinking about other solutions."
Well that's kind of the point of the article. That in order to "win", the US needs an incentive structure that encourages open models.
I'm not sure what that looks like though.
Why do people always bring up state support when it comes to China? As if the U.S. doesn't provide massive tax breaks and explicit funding to industry?
It's on every tech post about China, as if it gives them some sort of "unfair" advantage.
Because the scale of the subsidization is unimaginably different.
The US is trillions in debt pal. Talk about unfair advantages!
The scale of public spend is probably much higher in the US..
yep. the ppp loans alone dwarf any state subsidies any other country has ever done
How much is it?
Don't have a choice, will probably have to go open-weights models as currently, "AI" is gated using 'whatwg cartel' web engines.
In the light of this, I am mechanically a proponent of very good open weights models, which I can download (for instance on on bittorrent) and run, slowly (the price), on local hardware.
That would be for coding.
If china puts its AI models on the same ground than US capital investment funds and big tech financial support (aka Big Tech international finance), they will very probably lose everything (know how, ML and inference infrastructures).
This is precisely why we need more projects like this https://github.com/bigscience-workshop/petals
There's no fundamental reason why models couldn't be developed and trained using community efforts. It might not be as fast and efficient, but it's definitely possible.
See the recent development of DiLoCo at Nous Research and Prime Intellect.
I am also confused by this point. The American government could force OpenAI and Anthropic to open their models, but then they would instantly evaporate, right? It doesn't seem like a choice that they can make, so framing it as a "winning" strategy doesn't make any sense to me. In what world could those companies have existed and opened their models?
Why did Google give Android away for free?
But they could do that because it didn’t cost them hundreds of billions of dollars to create Android and they could retain 99% of control over end users.
Because Android was the door to there services. That's very much the same situation we have with model creation right now. China understood it, they act on it and as a result they are winning.
They could, but they don’t have to. The Chinese have beaten them to it, and the rest of the world will benefit from it and the circle will be complete once the models get a little bit faster/smaller and the localized hardware does the same and it will, it is inevitable.
The one thing that is sort of ironic or bad is that between Russia and the Ukraine there’s a large number of mathematically inclined people that if it wasn’t for the Putin war, their brain power working on AI models would have probably pushed open source down the road, even faster…
> The Chinese have beaten them to it, and the rest of the world will benefit from it and the circle will be complete once the models get a little bit faster/smaller and the localized hardware does the same and it will, it is inevitable.
This reads just like "AGI is 2 years away", I'll go set my calendar...
But I still don't get it. Like China could be come the world's leading producer of chocolate...if they started giving away chocolate for free. Would we be having this conversation saying that Switzerland lost because they were greedy and protectionist and didn't decide to give away their chocolate for free first (I realize Switzerland probably isn't actually the world's top producer of chocolate).
Long history of open source projects already have answers on how to monetize a free and complex open source product.
- Low development cost: collaborative efforts from open source contributors, innovative model training and serving for llm (Chinese models costs a fraction to train and their local chip design and manufacturing are catching up, plus cheap electricity)
- monetizing by selling hosted services, while leaving the core product free to tinker with / self host. China’s gdp is 2/3 of the US and it’s already a huge market for AI - which OAI and A\ don’t enter.
- for (the US) market that they can’t enter, let the US cloud providers to do free marketing / advocacy for them. Gaining share of mind. It costs them nothing.
The idea that Chinese models cost less to train seems to be based on that one time DeepSeek estimated the training cost for their V3 model at GPU rental rates as $5 million, and comparing this to other companies' entire R&D budgets. Yet DeepSeek raised $7 billion of fresh money last month, enough to train more than 1000 such models. What gives?
- You need to train lots of experimental models to dial in the training process just right for the one model that actually gets released in the end. Fortunately, these can be smaller.
- However, everyone is training much bigger models now, and doing a lot of RL rollouts on top.
- You can't get the GPUs for this piecemeal at rental rates because they need to be wired together using high-bandwidth interconnects.
- Nvidia GPUs are much more expensive in China, and local alternatives are still immature and not as efficient. Some companies have gotten around this using data centers in Singapore, which should tell you that electricity prices are not the primary consideration.
- The one line item where Chinese companies can probably save quite a bit of money is salaries for rank-and-file researchers.
In any case, they need to make back that money somehow. Giving away freebies isn't going to cut it.
In Russian opposition's mostly liberal discussions their school of thought connects several things together (sorry for not going directly to Marx's "General Intellect" and "Fragment on Machines" and using AI summaries instead ) - general idea of communism in China vs. techno-libertarianism of Thiel, Musk and the likes, and the Marx's thinking like:
"Fragment on Machines":
"he explores how human knowledge and collective intellect become embedded into machines, divorcing the worker from their own creativity."
"General Intellect":
"These texts are widely discussed for his concept of the General Intellect—the idea that society's shared, collective knowledge increasingly drives production rather than raw manual labor, and that this knowledge is alienated from workers and used as an instrument of capital."
(note: my point isn't to pass any political judgement here, like what real communism in China or not real, is it good or bad, i just find it interesting that pure political discussions by people with no technical credentials bring AI as a major factor today)
> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.
Right... and there are two problems with this:
1. Eventually the capabilities of closed-weight models will just vastly outstrip open-weight models if the underlying assumptions about compute and scale needed are mostly on the mark. So you can release open-weight models and they will have great use cases and applications, but ultimately similar to how you don't use an open-source phone or a budget Android phone from Wal-Mart and you buy an iPhone instead, you will see that although they "do the same thing" one product is clearly superior and you just have to pay for it. For this to not be true...
2. then it incentivizes most (all?) companies, American, Chinese, or European to halt development of models because if you spend all the CAPEX and it can just be copied and turned open-source nobody will invest in that. Given that China is not halting development of proprietary models I believe the current strategy and the subsequent approach to release open-weight models is at best a stall tactic, and at worse a sign of desperation.
Open source and the support and development models around it have been great. But folks are a little too dogmatic about it. Open-source software isn't a moral good, and closed-source software isn't a moral wrong either.
Imagine there’s a school where all the kids there are being tutored by the best. Also imagine a bunch of neighboring schools drastically falling behind that would need insane amounts of money to keep up.
This becomes a problem because all the kids from the rich school will dominate the order schools. They’ll get even more money as time goes on from their kids paying it forward to the point where all other kids are bound to work for them.
Now let’s say one other school does have the money for best tutors, BUT they know they’ll run out pretty quickly. Instead of trying to compete in a losing game, they decide to give every school in the world access to their elite lesson plan. Now, for a time, everyone will be on close to a level playing field. If the other schools improve upon their own lesson plans and keep sharing them with others, one day the elite school will wake up to find they are no longer on top. The parents have started to move their kids to other schools because the rich school is no longer attractive at the high cost they charge students
Sure and to complete your analogy here, the rich schools realize that the curriculum they develop and put a lot of time and money into creating is just used by the cheap schools, so they stop developing it because nobody loses money for long and so neither the rich or cheap schools develop any new curriculum.
Now what?
The fundamental problem here is incentives and tactics. Either the models are actually better (which I think the iPhone to cheap Android phone really speaks to, i.e. they do the same thing but one is 50x better at 5x-10x the cost) and thus they can be gate kept and like the iPhone the vast majority of profits go to a select few with high end implementations. OR the models aren't actually that much better, companies lose a fortune and then nobody can create any better commercial models or build out scale needed for open source models because it's not profitable.
We could wind up with only open-source models or something along those lines, but if the compute and scale is needed to train the models, nobody will be able to do that profitably and so AI research is either gate kept and silo'd for something like military applications or it just doesn't really happen because there's no funding for this scale of build out.
Try selecting a file in your 50x iphone - triple copy of the same file or ateast double copy (assuming os will post a soft link). If a large file then you are toast.
Try mmapping > 5GB file in your 50x better iPhone.
Try running any service in the background.
The list goes on and on.
Your 50x better suddenly became 50x worse compared to a much cheaper android.
iphone is 50x better? because you get a blue box around your text rather than green?
androids and iphones are approximately the same thing
the kinda obvious direction LLM training can go is into the direction of particle physics, and the training is set up democratically and through universities and via multi-state funding
then the resulting weights end up open, the same as the particle detection data
> androids and iphones are approximately the same thing
Yet...
> and the training is set up democratically and through universities and via multi-state fundingPossible, certainly. But this case also applies to China and its "open-weights" strategy. They won't be able to form companies either or get ahead.
[1] https://www.digitalapplied.com/blog/mobile-os-market-share-2...
"Things on my phone cost more, so its clearly a better phone."
Based on the above stat, it sure seems like Androids more versatile and inexpensive for a much larger group of users.
Hmm...that sounds familiar....
Sure, but at the end of the day, Apple makes a lot more money not just on device sales but app sales too even though Android phones are allegedly more versatile and inexpensive.
You can talk about open-source and cheap Chinese models all you want, but at the end of the day if American companies are making all the money that's kind of all that matters. That will feed into development and maintaining an edge.
> so they stop developing it because nobody loses money for long and so neither the rich or cheap schools develop any new curriculum
Why do you assume the poor schools wouldn't be smart enough to keep it going? It's very likely the can collectively beat the rich school now that the one other rich school opened access to their materials and led the charge.
> but if the compute and scale is needed to train the models, nobody will be able to do that profitably
But they would. Efficiently hosting models will be the real business and early access to models with incremental improvements will not be the moat once thought. The reason other companies don't feel they can compete is the same reason OAI and Anthropic will lose their lead. They banked too heavily on another player NOT leading the charge on open research and poured disgusting amounts of money at closed source models.
China has proved they can take the limited resources available to them and build something better than what the US is offering consumers [1]. I'm just waiting for other countries to start pitching in.
Reminds me of the NSA and their early battles with cryptographers who believed in open research.
[1] https://x.com/DavidSacks/status/2078984980588531855
Software being open source has many strong positive externalities. It advances human knowledge and freedom. If you don't think that counts as a moral good then I'm baffled by what you think a moral good is.
It depends on how it is applied. You can release open source software that advances knowledge and freedom that results in economic destruction or the loss of life, for example.
Open source is in the tradition of humans sharing past knowledge, long-term we just can’t keep a secret it’s a time, honored tradition…
> Open-source software isn't a moral good
Yes it is.
Prove it
I’ll take a swing at it. Open Source is a form of sharing with the wider community. Closed Source is not sharing. Moral good is based on doing good outside of your own benefit (the opposite of selfishness.)
Ergo it’s a kind of moral good.
And I’m not even an advocate for open source.
This argument boils down to X is good, therefore more of X is good. But you can see how that breaks down with even trivial examples. Not a great argument.
The second piece of this "a moral good is based on doing good outside of your own benefit" - says who? Why? This logic is also faulty. You're also cargo-cutting self-interest in here as a moral failure when many good things depend on humans acting in their own self interest. For example I completely and selfishly installed a new tree at my house. But the community benefits from carbon capture, shade, &c.
I understand the sentiment you have here and I think for everyday use and having some guiding principles it is probably fine, but don't confuse this for a principle that is actually examined. You can find contradictions rather easily, never mind solid arguments which expose cases where what you think is true is not really true and so forth.
Thanks for the discussion.
>This argument boils down to X is good, therefore more of X is good.
No, I only argued that it was a moral good, the kind of good. I actually may disagree with others about whether you should pursue a good just because it’s good.
>says who? Why?
Good question, it’s just a common framing that I see in classical discussions. I didn’t intend for it to be exclusive, I think there’s moral good outside of that.
>don't confuse this for a principle that is actually examined
I hear you, I think this is a simplified version suitable for an online comment. In particular I’m not saying that if you do something other than a moral good then you are doing something wrong. There are many actions that are morally neutral. Also it is possible to construct artificial situations where you may violate some moral good in pursuit of another.
That's fair, I don't really take a large issue with anything you wrote here. My general point was to shake loose some dogmatic apples hanging on to their naive and unsubstantiated beliefs that open-source must be some moral good and closed-source must be some moral bad. Both can be good or bad.
Thanks
Justify your assertion and let people disprove it
No u
That's just like, your opinion.... man....
It's important also that open-weight isn't open source. If you can't download the training data (fully labeled), source code of the NN, and follow the README to build and train it yourself assuming oyu had the hardware then it's not open source.
tldr there's no "source" in open weight models therefore they are not open source.
Exactly. AFAIK none of the popular "open Chinese" models have published the full pre- and post-training pipeline, so the models are only partially open, if at all (plus the openly shared final weights, of course).
Did you just assume the singularity?
No