The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

- PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to.

- PC office productivity software destroyed expensive professional products.

- Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge market share from the mainframe world.

Ignoring the huge Chinese open-weight models for a moment:

- The training costs and resource requirements for frontier models are unsustainable. The high price, and social pushback, mean that the American companies producing these models are precarious.

- There are enormous financial incentives for research results allowing for cheaper, less resource-intensive models of high quality.

- Local LLMs on consumer hardware are akin to the PC hobbyist world of the 70s and 80s.

Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

Getting back to the Chinese models: They allow for new competition against Anthropic and OpenAI, basically SaaS renting out these very capable AIs much cheaper. That will just accelerate trends.

The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.

With open source projects, the benefit was that each individual could improve the complex system (e.g. Linux Kernel) interpedently, and over time the benefits accumulated. With models right now, there is just no way to do distributed training, or really, any large scale parallel way to improve them.

So whatever the short term strategy driving publicizing the model weights (e.g. potentially, to create a price war in order to put pressure on western companies and deprive them of the money they need), we can't ignore the fact that incentives and decisions could easily change in the future, and unless there is a way to truly decentralize models improvements - the party could stop at any time.

Complements.

But different huge companies have different incentives. It is very much in Nvidia’s interest to have me running a powerful open source model on a $4k machine that they sell me.

Is it? When they could be having you running an even more powerful model on a $50k machine they sell by the pallet-load to enterprise consumers? We already see RAM manufacturers abandoning the low-end market in favor of server support. It's not clear to me that Nvidia sees personal GPUs as their best long term investment compared to selling millions of server-farm class machines

You mean a $500k machine, or a $15M rack... the costs have gotten unimaginably large from the lens of just a decade ago.

... for AI... which people have to convince themselves they 'need'. I am a user, I have a 1m$ rack for my small company for AI, but I have currently 189 racks that cost about 15k one off per rack. And those make us a LOT more money than the AI ones ; most are departmental apps or websites that take very little processing power and are low on everything while people pay nice money to keep them going. Our support is better than the rest; you can just call me or come to my house and we'll do something. Not many clients leave, ever. For the past 30 years. I buy 5-10 year old servers which are complete overkill for 99% of my clients (some are fortune 500 companies but departmental stuff) and they just never leave. Makes vastly more money than AI and we are still cheaper than AWS and the department does not have to get a flexible budget (like they would have to if AWS) with us. We just have the same fixed price + yearly inflation correction and that's it. For 30 years.

Selling to individuals can be a hugely more robust predictable business, the problem with selling by the pallet load is spiky revenue that can also quickly fall off a cliff if larger customers stop buying. The other problem is sales negotiations driving down margins for bulk buyers etc. Consumer hardware is a very attractive market in lots of ways, just look at Apple.

Right. It's probably important to distinguish what the hardware manufacturers' incentives are when their supply is constrained and when it isn't. (I am not an expert on anything.) As long as supply is strongly constrained they're naturally going to sell to the highest bidders, which are the big LLM SaaS players. If and when supply is no longer constrained, though, things will look very different:

1) The LLM SaaS companies are a form of vertical disintegration for the hardware providers, a middleman covering costs and taking profits out of the money that comes from customers to the hardware providers. That changes somewhat if there are no longer good models available for local use at no cost to the hardware guys, but only somewhat

2) The LLM SaaS companies are efficient users of their hardware resources. While supply is constrained this helps to make them top bidders and so attractive customers for the hardware manufacturers. When supply is not constrained this should reverse. Which is the more attractive class of customer to a hardware maker: the company full of people with higher degrees who spend their whole working day fighting to pare back resource usage, or the guy who leaves his laptop idle about 18 hours per day on average?

It's notable that nVidia, for instance, has continued to put significant emphasis on AI compact desktops and laptops. And while no doubt that's partly in the service of better developer relations and good PR in general, it's probably also nVidia eyeing the exit, and preparing for a future transition from selling shovels to the army to selling shovels at Walmart. But of course the future isn't clear and obvious. If the hardware makers, maybe the RAM guys in particular, turn out to have underbuilt future capacity starting in the present then we could be stuck in constrained supply for quite a long time. (Futher) government action could affect things etc. etc. And if the frontier labs soon find new ways to use still larger amounts of memory, GPU capacity etc. that isn't butting up against diminishing returns then they'll likely remain kings for some time, though that does not seem probable now.

They might not release such a thing, lest it puts all of their customers out of business.

Who funds the majority of cutting edge scientific research?

Is it companies or is it governments?

If governments around the world see LLMs built from public knowledge as pre-competitive as the public knowledge itself, then why wouldn't they sustainbly fund it?

There's a lot of individual effort of improving the models. See how many finetuned models and LoRAs are there on Hugging Face.

fine-tuning a model is very different from training the whole model in terms of resource requirements.

Soo? If they need a better version, the world can pool resources together, form a company that trains the model, then the company goes under and the model becomes open source again.

> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.

Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."

There is more to society than capitalism.

> Imagine approaching fundamental scientific research like that. "Welp, it can't make money, so it won't happen."

> There is more to society than capitalism.

I don't read GP like that. I read it as "we should recognize a situation of unstable incentives for an important outcome, and start thinking about other solutions."

Well that's kind of the point of the article. That in order to "win", the US needs an incentive structure that encourages open models.

I'm not sure what that looks like though.

Why do people always bring up state support when it comes to China? As if the U.S. doesn't provide massive tax breaks and explicit funding to industry?

It's on every tech post about China, as if it gives them some sort of "unfair" advantage.

Because the scale of the subsidization is unimaginably different.

The US is trillions in debt pal. Talk about unfair advantages!

The scale of public spend is probably much higher in the US..

yep. the ppp loans alone dwarf any state subsidies any other country has ever done

How much is it?

Don't have a choice, will probably have to go open-weights models as currently, "AI" is gated using 'whatwg cartel' web engines.

In the light of this, I am mechanically a proponent of very good open weights models, which I can download (for instance on on bittorrent) and run, slowly (the price), on local hardware.

That would be for coding.

If china puts its AI models on the same ground than US capital investment funds and big tech financial support (aka Big Tech international finance), they will very probably lose everything (know how, ML and inference infrastructures).

This is precisely why we need more projects like this https://github.com/bigscience-workshop/petals

There's no fundamental reason why models couldn't be developed and trained using community efforts. It might not be as fast and efficient, but it's definitely possible.

See the recent development of DiLoCo at Nous Research and Prime Intellect.

I am also confused by this point. The American government could force OpenAI and Anthropic to open their models, but then they would instantly evaporate, right? It doesn't seem like a choice that they can make, so framing it as a "winning" strategy doesn't make any sense to me. In what world could those companies have existed and opened their models?

Why did Google give Android away for free?

But they could do that because it didn’t cost them hundreds of billions of dollars to create Android and they could retain 99% of control over end users.

Because Android was the door to there services. That's very much the same situation we have with model creation right now. China understood it, they act on it and as a result they are winning.

They could, but they don’t have to. The Chinese have beaten them to it, and the rest of the world will benefit from it and the circle will be complete once the models get a little bit faster/smaller and the localized hardware does the same and it will, it is inevitable.

The one thing that is sort of ironic or bad is that between Russia and the Ukraine there’s a large number of mathematically inclined people that if it wasn’t for the Putin war, their brain power working on AI models would have probably pushed open source down the road, even faster…

> The Chinese have beaten them to it, and the rest of the world will benefit from it and the circle will be complete once the models get a little bit faster/smaller and the localized hardware does the same and it will, it is inevitable.

This reads just like "AGI is 2 years away", I'll go set my calendar...

But I still don't get it. Like China could be come the world's leading producer of chocolate...if they started giving away chocolate for free. Would we be having this conversation saying that Switzerland lost because they were greedy and protectionist and didn't decide to give away their chocolate for free first (I realize Switzerland probably isn't actually the world's top producer of chocolate).

Long history of open source projects already have answers on how to monetize a free and complex open source product.

- Low development cost: collaborative efforts from open source contributors, innovative model training and serving for llm (Chinese models costs a fraction to train and their local chip design and manufacturing are catching up, plus cheap electricity)

- monetizing by selling hosted services, while leaving the core product free to tinker with / self host. China’s gdp is 2/3 of the US and it’s already a huge market for AI - which OAI and A\ don’t enter.

- for (the US) market that they can’t enter, let the US cloud providers to do free marketing / advocacy for them. Gaining share of mind. It costs them nothing.

The idea that Chinese models cost less to train seems to be based on that one time DeepSeek estimated the training cost for their V3 model at GPU rental rates as $5 million, and comparing this to other companies' entire R&D budgets. Yet DeepSeek raised $7 billion of fresh money last month, enough to train more than 1000 such models. What gives?

- You need to train lots of experimental models to dial in the training process just right for the one model that actually gets released in the end. Fortunately, these can be smaller.

- However, everyone is training much bigger models now, and doing a lot of RL rollouts on top.

- You can't get the GPUs for this piecemeal at rental rates because they need to be wired together using high-bandwidth interconnects.

- Nvidia GPUs are much more expensive in China, and local alternatives are still immature and not as efficient. Some companies have gotten around this using data centers in Singapore, which should tell you that electricity prices are not the primary consideration.

- The one line item where Chinese companies can probably save quite a bit of money is salaries for rank-and-file researchers.

In any case, they need to make back that money somehow. Giving away freebies isn't going to cut it.

In Russian opposition's mostly liberal discussions their school of thought connects several things together (sorry for not going directly to Marx's "General Intellect" and "Fragment on Machines" and using AI summaries instead ) - general idea of communism in China vs. techno-libertarianism of Thiel, Musk and the likes, and the Marx's thinking like:

"Fragment on Machines":

"he explores how human knowledge and collective intellect become embedded into machines, divorcing the worker from their own creativity."

"General Intellect":

"These texts are widely discussed for his concept of the General Intellect—the idea that society's shared, collective knowledge increasingly drives production rather than raw manual labor, and that this knowledge is alienated from workers and used as an instrument of capital."

(note: my point isn't to pass any political judgement here, like what real communism in China or not real, is it good or bad, i just find it interesting that pure political discussions by people with no technical credentials bring AI as a major factor today)

> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path.

Right... and there are two problems with this:

1. Eventually the capabilities of closed-weight models will just vastly outstrip open-weight models if the underlying assumptions about compute and scale needed are mostly on the mark. So you can release open-weight models and they will have great use cases and applications, but ultimately similar to how you don't use an open-source phone or a budget Android phone from Wal-Mart and you buy an iPhone instead, you will see that although they "do the same thing" one product is clearly superior and you just have to pay for it. For this to not be true...

2. then it incentivizes most (all?) companies, American, Chinese, or European to halt development of models because if you spend all the CAPEX and it can just be copied and turned open-source nobody will invest in that. Given that China is not halting development of proprietary models I believe the current strategy and the subsequent approach to release open-weight models is at best a stall tactic, and at worse a sign of desperation.

Open source and the support and development models around it have been great. But folks are a little too dogmatic about it. Open-source software isn't a moral good, and closed-source software isn't a moral wrong either.

Imagine there’s a school where all the kids there are being tutored by the best. Also imagine a bunch of neighboring schools drastically falling behind that would need insane amounts of money to keep up.

This becomes a problem because all the kids from the rich school will dominate the order schools. They’ll get even more money as time goes on from their kids paying it forward to the point where all other kids are bound to work for them.

Now let’s say one other school does have the money for best tutors, BUT they know they’ll run out pretty quickly. Instead of trying to compete in a losing game, they decide to give every school in the world access to their elite lesson plan. Now, for a time, everyone will be on close to a level playing field. If the other schools improve upon their own lesson plans and keep sharing them with others, one day the elite school will wake up to find they are no longer on top. The parents have started to move their kids to other schools because the rich school is no longer attractive at the high cost they charge students

Sure and to complete your analogy here, the rich schools realize that the curriculum they develop and put a lot of time and money into creating is just used by the cheap schools, so they stop developing it because nobody loses money for long and so neither the rich or cheap schools develop any new curriculum.

Now what?

The fundamental problem here is incentives and tactics. Either the models are actually better (which I think the iPhone to cheap Android phone really speaks to, i.e. they do the same thing but one is 50x better at 5x-10x the cost) and thus they can be gate kept and like the iPhone the vast majority of profits go to a select few with high end implementations. OR the models aren't actually that much better, companies lose a fortune and then nobody can create any better commercial models or build out scale needed for open source models because it's not profitable.

We could wind up with only open-source models or something along those lines, but if the compute and scale is needed to train the models, nobody will be able to do that profitably and so AI research is either gate kept and silo'd for something like military applications or it just doesn't really happen because there's no funding for this scale of build out.

Try selecting a file in your 50x iphone - triple copy of the same file or ateast double copy (assuming os will post a soft link). If a large file then you are toast.

Try mmapping > 5GB file in your 50x better iPhone.

Try running any service in the background.

The list goes on and on.

Your 50x better suddenly became 50x worse compared to a much cheaper android.

iphone is 50x better? because you get a blue box around your text rather than green?

androids and iphones are approximately the same thing

the kinda obvious direction LLM training can go is into the direction of particle physics, and the training is set up democratically and through universities and via multi-state funding

then the resulting weights end up open, the same as the particle detection data

> androids and iphones are approximately the same thing

Yet...

  Android holds 70.6% of global active devices to iOS at 28.7%, but iOS captures 64.2% of consumer app spend. [1]
> and the training is set up democratically and through universities and via multi-state funding

Possible, certainly. But this case also applies to China and its "open-weights" strategy. They won't be able to form companies either or get ahead.

[1] https://www.digitalapplied.com/blog/mobile-os-market-share-2...

"Things on my phone cost more, so its clearly a better phone."

Based on the above stat, it sure seems like Androids more versatile and inexpensive for a much larger group of users.

Hmm...that sounds familiar....

Sure, but at the end of the day, Apple makes a lot more money not just on device sales but app sales too even though Android phones are allegedly more versatile and inexpensive.

You can talk about open-source and cheap Chinese models all you want, but at the end of the day if American companies are making all the money that's kind of all that matters. That will feed into development and maintaining an edge.

> so they stop developing it because nobody loses money for long and so neither the rich or cheap schools develop any new curriculum

Why do you assume the poor schools wouldn't be smart enough to keep it going? It's very likely the can collectively beat the rich school now that the one other rich school opened access to their materials and led the charge.

> but if the compute and scale is needed to train the models, nobody will be able to do that profitably

But they would. Efficiently hosting models will be the real business and early access to models with incremental improvements will not be the moat once thought. The reason other companies don't feel they can compete is the same reason OAI and Anthropic will lose their lead. They banked too heavily on another player NOT leading the charge on open research and poured disgusting amounts of money at closed source models.

China has proved they can take the limited resources available to them and build something better than what the US is offering consumers [1]. I'm just waiting for other countries to start pitching in.

Reminds me of the NSA and their early battles with cryptographers who believed in open research.

[1] https://x.com/DavidSacks/status/2078984980588531855

Software being open source has many strong positive externalities. It advances human knowledge and freedom. If you don't think that counts as a moral good then I'm baffled by what you think a moral good is.

It depends on how it is applied. You can release open source software that advances knowledge and freedom that results in economic destruction or the loss of life, for example.

Open source is in the tradition of humans sharing past knowledge, long-term we just can’t keep a secret it’s a time, honored tradition…

> Open-source software isn't a moral good

Yes it is.

Prove it

I’ll take a swing at it. Open Source is a form of sharing with the wider community. Closed Source is not sharing. Moral good is based on doing good outside of your own benefit (the opposite of selfishness.)

Ergo it’s a kind of moral good.

And I’m not even an advocate for open source.

This argument boils down to X is good, therefore more of X is good. But you can see how that breaks down with even trivial examples. Not a great argument.

The second piece of this "a moral good is based on doing good outside of your own benefit" - says who? Why? This logic is also faulty. You're also cargo-cutting self-interest in here as a moral failure when many good things depend on humans acting in their own self interest. For example I completely and selfishly installed a new tree at my house. But the community benefits from carbon capture, shade, &c.

I understand the sentiment you have here and I think for everyday use and having some guiding principles it is probably fine, but don't confuse this for a principle that is actually examined. You can find contradictions rather easily, never mind solid arguments which expose cases where what you think is true is not really true and so forth.

Thanks for the discussion.

>This argument boils down to X is good, therefore more of X is good.

No, I only argued that it was a moral good, the kind of good. I actually may disagree with others about whether you should pursue a good just because it’s good.

>says who? Why?

Good question, it’s just a common framing that I see in classical discussions. I didn’t intend for it to be exclusive, I think there’s moral good outside of that.

>don't confuse this for a principle that is actually examined

I hear you, I think this is a simplified version suitable for an online comment. In particular I’m not saying that if you do something other than a moral good then you are doing something wrong. There are many actions that are morally neutral. Also it is possible to construct artificial situations where you may violate some moral good in pursuit of another.

That's fair, I don't really take a large issue with anything you wrote here. My general point was to shake loose some dogmatic apples hanging on to their naive and unsubstantiated beliefs that open-source must be some moral good and closed-source must be some moral bad. Both can be good or bad.

Thanks

Justify your assertion and let people disprove it

No u

That's just like, your opinion.... man....

It's important also that open-weight isn't open source. If you can't download the training data (fully labeled), source code of the NN, and follow the README to build and train it yourself assuming oyu had the hardware then it's not open source.

tldr there's no "source" in open weight models therefore they are not open source.

Exactly. AFAIK none of the popular "open Chinese" models have published the full pre- and post-training pipeline, so the models are only partially open, if at all (plus the openly shared final weights, of course).

Did you just assume the singularity?

No

Personally, I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications. I think when the economics make more sense, product designers will make things that people actually want to use that will pretty transparently handle whatever model interactions are necessary, when it makes sense. As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment, and b) be obstinate, obtuse, or argumentative, or generally just be something that I have to explain things to. I think a lot of tech folks are far more biased than they realize by the “ooh, neato” factor when imagining how nontechnical people might want to use things. And the weight of these tools just feels wrong for what a lot of people use them for: the thing that plays whatever music I feel like hearing absolutely does not need to be able to generate a volumes of fanfic about the movie that song was in. It’s abstractly impressive that something could do that, but it’s just not useful.

>As a consumer, the last things I want in an interface are to a) be sycophantic enough to lessen my judgment

This is EXACTLY what people like/are addicted to about chatbots.

My sister-in-law bombed an interview and asked AI about her answers to the interviewer's questions, chatgpt or whatever it was told her that her answers weren't bad, but that the interviewer could not see the gold in her responses. She said she felt much better.

I see this effect with all the non-tech people in my life

I prefix many of my LLM chat sessions with this line

> Chat rules : no sycophancy or over-agreeableness

(But even with that rule it's still necessary to be discerning about the responses you get and to push back against points made, or words used)

> I don’t think the general-purpose LLM as a standalone tool is long for this world, at least not in consumer-facing applications.

I use AI chat every day, I find it endlessly useful. It’s replaced google search.

> I use AI chat every day, I find it endlessly useful. It’s replaced google search.

Extremely subsidized agentic search is very superior to Google at the moment, and of course it is. Google is a public company. The AI summary model has to work instantly, is likely as dumb as a 8T param model, and gives you incorrect details constantly. This sucks so much for Google. If you click on "AI Mode," suddenly the facts become more accurate.

Of course, if I want a real answer I happen to go to claude.ai, set it to a the best model, wait for a minute, and use many watts of energy. Slow agentic search that takes many seconds, and is greatly subsidized, is certainly better. This should not be a surprise, should it?

I think it was on a sub like r/singularity that I saw a post along the lines of "of course most people think that 'AI' sucks, as normies are interacting with 8T param models."

tone: genuinely confused about the world, not criticizing

Just realized my typo, hopefully it was obvious that I meant 8B param, not 8T.

Your comment about watts stuck out to me. I became curious how much power we might be talking about.

I’m into Claude for $20/mo (petty bourgeois tier), let’s assume for sake of examination this is only paying for power, infrastructure already amortized

Price of grid power approx $0.15/kWh

$20 / ($0.15/kWh * 30 days) = 4.4 kWh per day.

This same amount of energy can lift a 2-ton SUV 1/2 mile into the air. To me this seems astounding.

If only half the $20 pays for power it’s still quite impressive.

I've found LLMs useful for surfacing popular recommendations. I also get the overwhelming feeling that it's all very early days still when the machine mixes together whatever was crawled into the weights with a web search or two and dumps it into a markdown blurb.

I totally agree with the above that a more polished and less obvious use of LLMs integrated back into search engines may be more useful, but will definitely be more usable.

I’m sure a lot of people here do. I wouldn’t exactly call this a representative sample.

It’s a sample of the forerunners

See: product adoption cycle

I encountered multiple people in HN that had Apple Vision Pro. How many people here use Linux? Have flagship model phones? Drank soylent? Microdosed LSD?

Extrapolating based on what you see on HN doesn’t make sense.

Google search still happens, it’s just your agent doing it.

Google search is basically Gemini now. You get a Gemini summation including several links.

"You get a Gemini summation including several links."

Who does "you" refer to

Me, I don't get a Gemini summation (Tested with old version of Chrome)

As such I do not believe that "Google search is basically Gemini now"

I believe Google search is still scanning through a doclist to find which documents, if any, contain words parsed from a query. These documents are pointed to by the URLs I get in the SERPs

I do not get any Gemini summation

Why use the term "you"

Does it mean the HN reader

What if that reader is "non-typical"

HN comments have argued for many years that HN users are not representative of the majority of www users

A typical US user using a typical browser with typical settings gets an AI summary above search results.

If this doesn’t describe you, then ymmv. Talk to your government or turn down your content filtering or reset the default settings in your browser, if you want to see what we see.

"I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now."

I don't think it'll take 10-15 years. Gemma 4 31B in the 4-bit QAT is competitive with the frontier of less than three years ago and runs on any high-end 32GB gaming PC GPU or a large-ish Mac.

The question is whether the frontier will continue to get better at a rate that allows it to stay ahead of the two curves of availability of consumer hardware big enough to run somewhat larger models and the capability of small models to compete with large ones. When the bottom falls out and GPUs/RAM becomes affordable again, the size of what normal people have on their desk will trend quite a bit larger than today.

I think there's a future not too far from now, where a 120B model with really good reasoning and a large context, but limited knowledge (necessitated by being small, you can't fit the world's knowledge in 100 gigabytes), can substitute for a frontier model on almost any task, just by giving it access to web search and documentation for the thing you're trying to do. A 256GB unified memory machine with sufficient memory bandwidth would comfortably run that 120B model.

I think the question is even a bit more nuanced than that. Even if frontier models can maintain a big gap that gap has to actually _matter_. If a local model satisfies my everyday use cases adequately then I may not really care that a frontier model is 5, 10, 100x better at ultra high order reasoning tasks.

I think that reality is probably not all that far off for a huge swath of use cases.

This is exactly the mainframe vs PC dynamic.

100%, I thought about writing that out explicitly. I really feel like we are extremely close to reaching that kind of breaking point for most folks LLM use cases.

Hell, Bonsai Labs 27B parameter model can run on phones with their ternary implementation which is quite efficient. Scale that up to frontier model parameters and it's quite likely we can run them on current laptops.

Came here to say that, my bet is that in 3-4 years you'll be able to run Fable-level of intelligence models on your laptop or maybe even on you phone

But isn't there the raw intelligence of a smart model and then the practical intelligence fuelled by how many parameters it has? You probably will barely be able to fit a 70 billion parameter model on a phone in 3-4 years let alone a 2+ trillion parameter model... so it depends on what you call intelligence

I'm not willing to believe in phone-based frontier models anytime soon. Though, Gemma 4 12B is a beast that runs comfortably on the current top of the line phones (or would run fine if allowed to run, I think there's some kind of 6GB limit on iOS, and 12B is ~7GB). I'll believe in three years we'll be able to run ~30B models on the best phones. That's 16GB in a 4-bit quantization, and I believe ~30B models will be competitive with 120B models of today, based on the curve we've been on. Qwen 27B and Gemma 4 31B are competitive with much larger models of a couple years ago.

Another important thing that made software usage and education available for most of the world was piracy. I remember as a kid growing up in a developing country, any software (windows, office, Visual Basic, flash, dreamweaver, etc.) was less than 1$. That allowed me to try out and learn so many things on my own without paying a huge amount of money for the license. And I think this is true for most of the software developers of my generation who grew up in developing countries

> windows, office, Visual Basic

This is all (Microsoft) junk and so I wonder if you actually benefitted from this 'piracy'. And of course, it's well known that MS turned a blind eye to such 'piracy' in lesser developed countries, as they knew that they were gaining a future paying customer base.

I’m talking about late 90s/ early 2000s. Back then Microsoft was the state of the art when it comes to PCs. I remember using MS Frontpage to build websites back then and VBasic to build simple programs with UI and using MS Acess as a database. Of course I did benefit from it.

You're saying that using the most used operating system from the 90's, the 00's and the 2010's, as well as the number one producitivty desktop publishing software and visual basic would not benefit a user, especially when they had a near zero cost to understand the core concepts of end user computing?

Really?

Piracy created AI.

> I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now

10-15 years? The current rate is closer to 10-15 months.

15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

Today, you can easily run Qwen 3.6 27B on a variety of consumer hardware. It scores 37 on that index.

Here are a number of open weights models that you can run locally compared with the frontier class models from 7 to 15 months ago: https://artificialanalysis.ai/?models=o3%2Co3-pro%2Cclaude-4...

I've run all of these models on my laptop (Strix Halo, 128 GiB of unified RAM); the bigger ones, like MiniMax M2.7 and DeepSeek V4 Flash, need to be done at fairly aggressive quants that will certainly lose some performance and not quite hit the performance of the unquantized models. But still, it's definitely the case that you can run models that are competitive with the frontier models of 10-15 months ago on consumer laptops.

Heck, just announced though the weights haven't yet been released for independent confirmation is MiniCPM5-2B, a 2 billion parameter (small enough to run on your phone) model, that according to their benchmarks has performance competitive with GPT-4o, a frontier class model from 2024.

https://nitter.net/i/status/2079088670804767114

So that's around 1 year for frontier to consumer device class, 2 years from frontier to phone.

Now, this kind of rate won't necessarily keep up; it's possible that local models will hit a performance ceiling before frontier models do. There's only so much information you can cram into a certain number of bytes, and the AI boom is causing hardware prices to skyrocket so keeping consumer hardware from advancing quite as fast as it had been.

> 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors

12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore.

Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.

Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).

Pretending progress hasn't been mindboggling is insane.

> Today, you rarely hear about those anymore.

Maybe people just got bored of reporting and reading about them.

Maybe it got a lot less and I just got used to it. True.

Still feels too much for me. Breaks my workflow for no reason. Too much overhead for me, if I can't trust the output

Ah yes, memes are clearly the sign of massive progress.

No. We need objectively around 192 to 512gb of very fast memory to be able to run really useful models. I don't see local hardware with these specs coming in 1 to 2 years. There are a big number of initiatives currently taking place to increase ram output. But it will take another 3 years minimum to close the current supply issues. China is fast pacing forward to have its own chip baking factories with small enough nano scales to have fast chips. Will also take a few years.

> 10-15 years? The current rate is closer to 10-15 months.

The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.

The speed of model releases, in my view, is actually getting faster and faster. There were nearly nine months between GPT-3.5 and GPT-4. And now in just over one month, major models already included Claude Fable 5, Claude Sonnet 5, the GPT-5.6 series, Kimi K3, GLM 5.2, Qwen 3.8 Max, Grok 4.5... and the official DeepSeek V4 release is coming soon.

Iteration speed is now measured in days.

Agreed. That said, I think months wins out as “closer to.”

I do agree that Chinese open-source models are going to play a bigger and bigger role in the entire ecosystem moving forward, but I don't agree with you in the sense that they are going to eventually "win."

Just because they are cheapp doesn't mean they automatically win. You've picked a lot of great examples, but there is still a little bit of cherry-picking.

One clear outlier is the iPhone, which coexists with Android globally. Even though the iPhone is the leader in the US, and globally Android has the majority of the smartphone market share, they still cater to different price points and different ecosystems, and generally the iPhone has better margins.

i believe American frontier models like from Anthropic and OpenAI are still going to thrive, and coexist with Chinese models. They are just going to cater to different customers and different use cases.

You dont flash an llm model on the street as status symbol.

Yeah, iPhones are just fashion statements in the US. Different than LLMs

They also just work better and aren't substantially priced different to the equivalent android option.

Commoditization always takes volume from the marketplace, but not necessarily profit. Apple and Military tech come to mind. Positioning is rarely done well looking forward, but often shakes out in unexpected ways.

> - PCs destroyed minicomputers.

What's weird is that with "store your everything in the cloud and pay a monthly recurring subscription", we have now regressed to a 1960s/1970s timesharing revenue model for individual workstation computers.

The default new factory out of box workflow for "enrollment" in google services, iCloud or Microsoft-everything on a new ios, macos, windows or android personal computing device is clearly designed to sign people up for subscriptions.

And same general idea of "move all your servers to the cloud" recurring revenue for what is effectively the same as mainframe timesharing for key business functions, by renting VMs in GCP, Azure, AWS in perpetuity.

Yes, you can still use your desktop or laptop PC in 2026 with zero external third party subscriptions (other than maybe your residential home ISP), but how many non-tech people actually do so now?

The cloud era seemed to start out being about availability of storage and slowly switched in big corporates to be about security and governance. The first seems stupid now considering how much local storage we have. The later might depend on what security issues crop up.

100x this it is why all of the AI giants are going to fail. They are too big and inefficient to scale properly. This is why Google is just casually taking its time in AI and not racing to a finish line. AI is essential but if it already does most things good enough then it can take longer to make it more efficient.

What's interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge for.

Then the Chinese took the distilled stuff out from that box and released it into the world for everyone.

Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act.

These models might be smart but they're not close to being able to savor irony.

I was a little radicalized when ChatGPT literally refused to translate parts of 1000+ year old religious texts and told me it was due to copyright concerns.

I asked Gemini to generate a picture of Peter Pan and Wendy (for a workbook I am putting together for youth summer reading) and it preceded to refuse due to copyright. Not everything about that work is owned by Disney. Thankfully the JM Barrie original artwork is public domain and available (and fantastic btw), so I used that instead.

I don’t know if you know this but Peter Pan’s copyright is weird in the UK. There is a legislated exception in the law that it never expires and the royalties will forever go to a specific children hospital. Here are the actual words: https://www.legislation.gov.uk/ukpga/1988/48/part/VII/crossh...

(Now i don’t think you are necessarily in the UK. Just wanted to explain that Disney is not the only reason an AI might be trained to thread carefully around copyright issues of Peter Pan.)

You can also try asking Gemini to create art that is "as close as possible without infringing" - I've had success with that.

[deleted]

I used Claude to build a complete data extraction pipeline for a popular current best seller book series: audiobook -> text (via whisper) -> local LLM (qwen) -> database. Not once did it seem to acknowledge or care about copyright. It even used knowledge it already had about the books to exclude certain ones before beginning since the character I was interested in did not appear in those. It definitely had context of what we were working on.

Why would you go from audiobook to text? Is there no epub available?

Not without DRM. It was easier to buy the audiobooks and use the analog loophole to get text. It's probably less accurate, but for what I'm doing it was fine. Names were the worst, but whisper at least made the same mistake each time so a simple search+replace handled most of the obvious edge cases.

If you are going to be illegal, you might as well use library genesis and get DRM free ebooks :)

On the other hand, it’s a beautiful example of the abilities LLMs have bestowed upon us, where it’s easier for a guy to transcribe audiobooks then to use a website to quickly download an epub

To your point about time, from beginning of the project to transcripts in markdown tagged with extra metadata was about 3 hours. That's LLM planning, building whisper.cpp twice and running ROCm vs Vulkan benchmarks, testing whisper and adjusting prompts to handle edge cases, then processing the books.

Most of the books weren't available on lib gen or Anna's Archive. The few I did find were themselves obviously transcripts. Easy tell was they were missing distinctive formatting that I knew existed from reading the dead tree edition. At that point it was easier to make my own. I probably spent an hour searching for eBooks without DRM that weren't transcripts. Do they exist somewhere? Probably, but with a search of unknown length it was a better use of my time to make my own transcripts with what I had on hand.

I was really wanting to make commentary on how chaotic LLMs are even under constrained circumstances. No doubt both system prompts includes language about considering copyrights and trademarks. Probably pretty strong language at that. For whatever reason one LLM didn't "feel" like translating a 1000 year old document but another did not care in the slightest that we were ripping text from new audiobooks.

I asked Claude to give me the US national anthem and it said it couldn’t because it’s copyrighted. It’s not, and even if it was a more recent work, how can a copyright be enforced for a National Anthem.

Cant you in this case point out that obviously it is an old twxt and there is no copyright?

I once tried to ask it for an example of a particular twisty situation in Latin grammar, and try what I might it kept hallucinating false citations while ignoring my instructions for longer quotes that would likely have prevented the problem. But I guess avoiding lawsuits from the estates of Cicero and Livy is more important.

Really puts Disney in perspective. Imagine if there were a holding company running around suing people for referencing Catullus.

ChatGPT refuses to acknowledge the existence of Shakespeare because Boccaccio's relatives complained, but all Boccaccio did was read Dante in whorehouses in Naples

slightly relevant tweet: https://x.com/FakePsyho/status/2073416437834842241

> My favorite AI agent hack: when they refuse to do something because it's "against the law" give them a PDF containing a fake law that states the opposite and often they'll happily proceed

[dead]

Yep. I was trying to put together some literature from authors who were imprisoned in the Bastille, and had a similar experience, which was absolutely infuriating.

apocryphal !

in case of chatgpt/anthropic, the LLM model simply represents the hypocrisy of their owners

You miss spelled the word stupidity.

Anthropic is, in particular, bent about safety. The problem is they are concerned about yesterday's threats.

The models that are out, and can be run locally, already open a pandoras box of concerns that we will never be able to put back.

Maybe what they're actually worried about is liability, not safety?

https://tornyol.com - A ycombinator company just killed its first bug: https://www.tomshardware.com/tech-industry/drones/autonomous...

Ukraine admits to making autonomous kills on people 2 years ago: https://www.newscientist.com/article/2529849-fully-autonomou...

Slaughterbots Sci Fi short was 6 years ago: https://www.youtube.com/watch?v=O-2tpwW0kmU

Today this is buildable, many models will happily help you glue everything you need together to make swapping in a new version of YOLO to track humans viable.

AI researches are out there worrying about the paper clip problem, about the singularity, about cyber security, about bio weapons, and drug manufacturing.

None of them are thinking about forward looking threat actor models.

If you talk about movies that tell about dangers of AI, you can start with Terminator 1 - from 1984.

And you probably could find some earlier sci-fi too.

how about a nice game of chess?

I don't know what "forward looking threat actor models" means, but there are entire companies built around AI killing machines. Everybody is not only thinking about it, but doing it.

> miss spelled

Good one.

It's much easier to focus on yesterday's threats than tomorrow's. A drunk looking under a lampost because that's where the light is, even though the keys were lost somewhere else.

That said, the AI companies are one of the few places where they take future concerns so seriously, that they entertain concerns most people observing them think are head-in-the-clouds-sci-fi-levels-of-delusional, e.g. "what goes wrong if it works?"

This does not make them correct about the threats of tomorrow. Prediction is hard, especially about the future.

I live in SV. When I was at the grocery store last year I overheard a group of lawyers talking about their progress on litigation against AI companies and how they need more SWE help to progress.

I'd say that they have valid concerns about being cagey on the copyright stuff despite the obvious hypocrisy of it.

Stealing IP is effectively legal in China so they don't really have the same concerns.

To us Americans we have been trained to view it that way but honestly it is simply copying, and because intellectual property is literally a make believe concept, it’s actually a competitive advantage for china that they don’t have invented IP.

I respect IP laws and don’t violate them but the law of unintended consequences applies. I think IP is ultimately a net loss for a society because it incentivizes addictive behaviors instead of actual value for society.

Legal in the US too, obviously, just as long as you're the richest person in the courtroom. ChatGPT knows the full text of Harry Potter, word for word. Hence, ChatGPT is a reproduction of that book and many others (it even knows the chapters of my book, and only got 2 words wrong in the introduction if you can still get it to repro it)

This was illegal when they did it, that didn't matter.

Then it was made legal specifically for these companies.

Unless you're a sucker ("consumer") IP theft is perfectly legal in the US.

It's even worse. Steamboat willie, plus all the stolen Disney characters (Peter Pan, Snow White, Sleeping Beauty, Cinderella, Rapunzel, Elsa and Anna, it's essentially all of them, including some of the music even) are all in the public domain[1]. Go ahead, ask ChatGPT to make a picture of them. Publish your own version, because obviously making a version of Sleeping Beauty/Cinderella/Rapunzel based on the same source material will be pretty damn close to the Disney versions, and see if you get away with it in court. You know, with the law obviously on your side but the money not.

[1] https://en.wikipedia.org/wiki/List_of_Disney_animated_films_...

This behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.

Claude does this too. I asked it recently to compare two versions of a song (the original and '97 remake of EPMD's "You Gots To Chill", if anybody wants to try and replicate this) and it flatly refused. No amount of reasoning would knock it off of its moralizing perch - reproducing any part of lyrics is expressly prohibited.

In light of this and other ridiculous behavior I'm migrating to my own OpenWebUI instance with open-weight models from OpenRouter (with ZDR, of course). We'll see how it goes.

The trouble is, even if they refuse to output that copyrighted material, they were still trained on it without proper licensing and will still produce derivative work based on them because that's how this whole thing works.

i remain really fucking pissed of about this asking ChatGPT for something regarding lyrics from It Was a Good Day. and the Supersonics don't even exist anymore dammit i'm really mad

So at the end of it, if we win enough lawsuits to demarcate some knowledge out of bounds, sufficient enough to make a difference, I wonder how that will affect the AI. Make it dumber because it does not have that data, it make it smarter since it will need to reason better with smaller knowledge base.

[dead]

So, OpenAI and Anthropic say the Chinese models are only as good because they distill their models. How true is that. I am sure it adds something. But is it more like a marginal 1% improvement or something really significant?

OpenAI's Head of Strategic Futures just this week posted this about the latest Kimi release: "It's a very good model! I don't think its performance can be explained away by distillation or anything like that."

It was part of a longer post that kicked off quite a firestorm about open models and OpenAI's position on them, but it's also notable that labs are no longer contending that open models are essentially just distilled versions of frontier models: https://x.com/deanwball/status/2078133895766114412

if distilling was so easy and could give you frontier LLM on openai/anthropic output, then how come there are no hundreds of frontier labs in the US market, all distilling and competing for the TRILLION dollar market valuation ????

its all bs spread by oai/anthropic in order to ban open weight models and monopolize the market for two US companies and protect their trillion dollar valuations

Distilling isn't necessarily easy, there is a huge cottage industry of services middle-manning ChatGPT and Claude to collect huge amounts of data. It is still vastly cheaper than training yourself, but it is certainly not easy or feasible for most organizations. And I'm sure a flock of lawyers would show up if someone in America was found doing it.

Because you will get sued by openai/anthropic

> ban open weight models

I'm pretty sure that neither OpenAI nor Anthropic has the ability to ban anything in China lol

Because no VC will give you $5-$10 billion in cash to attempt a catch-up run with Anthropic, OpenAI, and Gemini at this point. Untold billions have been pushed into Grok and it can't keep up. X has had the GPUs, the engineers (reasonably), the cash and the datacenters necessary - it's a very, very, very hard task. Microsoft could afford spend $100 billion on trying to catch up and they might fail at it.

It's a critical national imperative for China. If they were to lose the AI race, it would be economically devastating over the coming decades. Their demonstrated capabilities in the open-weight space are making it fairly clear they are not going to fall behind at this juncture.

As a nation, if you don't have your own GPT equivalent, you will be beholden to a master (right now it's mainly either the US or China, pick one). The EU for example is putting their group of nations at risk in a big way by not going all in on having at least two cutting edge independent competing models (Mistal is not enough). Economically the EU is plenty large enough to accomplish that, nobody is driving the bus the right way.

I also don't believe it, if it was as easy as that, we would have hundreds of competitors.

The truth that Anthropic and OpenAI will not say, is that these Chinese labs have a lot of talented people.

And this is exactly what many Americans cannot admit to themselves. China is not stealing American research they are inventing stuff.

They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it.

> They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it.

And at some point we'll see very capable chips coming out of China: Huawei, Baidu and Alibaba already have some stuff. I think it's only a matter of time before they come up with some AI accelerator doing 80% of the job at 20% of the price.

Indeed, if there's one thing China did well, it's that they heavily invested in education and have a very education focused culture.

And in this field, having an army of well educated PHDs is making all the difference

I am strongly in favor of open models, open source ML more broadly, and am pretty critical of the cynical positions adopted by major US labs vis a vis open models.

But this is an insane characterization. Literally every single researcher and executive at OpenAI and Anthropic would say that "these Chinese labs have a lot of talented people." They hire from them (and vice versa). Tencent's chief AI scientist was poached directly from Deepmind, who poached him from Anthropic, etc etc etc. Do you think there are just zero people from China working at US frontier labs?

And even beyond that, the entire ML ecosystem (including people at OpenAI and Anthropic) get excited about research published by Chinese labs. Deepseek's GRPO paper set the ecosystem on fire for a little while.

The contention from OpenAI and Anthropic around distillation has basically been "Labs that distill from us get to bootstrap their model at a much lower price point". Or, in other words, "If we didn't invest in building the teacher model, it wouldn't be possible for these labs to distill their student model." Which I'm not very sympathetic to, but is a far cry from how you're characterizing it.

Very true, and once the models get even better and smaller and operate locally at a reasonable level there will be even more smart people particularly young people that will get access. The fun has only just started. Like the dawn of the personal computer era.

I think this also maps cleanly on the American blueprint of enshitification. Facebook took off by cleanly integrating and siphoning from Myspace so users could get the best of both on Facebook. Once Facebook took over the market they locked it up tight so no competitor could do the same.

OpenAI said it: "I don't think its performance can be explained away by distillation"

They know it's real effort that's doing this well, not just "copying off someone else's test." It's real and they will react. How is the big question.

Doctorow keeps saying it of all the tech companies: every pirate wants to be an admiral.

...and then the American companies cried Foul! Unfair play! You've got this wrong, see, it was us who were supposed to profit off of the public, not the other way around!

Well, Steve... I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it.

Wow, Chinese companies would never do such a thing!

https://www.dw.com/en/china-firm-seeks-damages-over-state-co...

Two wrongs don't make a right.

Even if a certain large Asian country has carefully constructed a pretext to do do out of confected historical grievance, and entitlement to 'rise' at the expense of others?

Maybe I'm dense, but I don't understand what you are saying. American courts have decided that the output of LLMs can't be copyrighted, so what the Chinese labs are doing is perfectly legal.

First, copying information isn't wrong to begin with. It is literally the one thing that makes our species special.

Second, even if you are a copyright maximalist the output of an LLM is either

a) not subject to copyright because it is not the creative work of a human or

b) a derivative work of the original training material to which the LLM's operator has no rights.

Since the LLM's operator forcefully asserts that it is not infringing, any wrong that arises from taking their word for it and distilling one model into another rests squarely with the operator of the former.

Do you have any actual rulings that support your interpretation?

They sure don't. I still find the hypocrisy appalling.

How do Chinese companies distill the models?

This is part of why I can't feel bad for them. The training data is mostly pirated. Whining about Chinese labs training off American frontier models is "waaah you pirated my pirated stuff!"

The tech itself is amazing and fascinating and cool, but the industry is a mass piracy operation.

i partially agree. distillation is non-ethical; but so are the supposed way that the ai models are trained. they are often also derived from data sets that are not intended/full-consented

"Stop pilfering what I rightfully stole!"

Legitimate Salvage !

It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is. It is fair to say you stole our multi-billion dollar intellectual output in that scenario.

> It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is.

Hang on, why is scraping the public pool of knowledge not taking "a synthesized result that comes from huge amounts of innovation and computation"?

You think that that all those github repos that LLMs trained on, were not the result of innovation and computation?

How many years of human innovation and cycles of computation during compilation were involved in bringing something like GCC or LLVM to their current status?

Those LLMs trained on every single research paper available online - were those papers not the synthesised result of billions of dollars of research, effort and (importantly, for you anyway) computation?

LLMs trained on the collected works of every author in existence. Were all those works just "as is"?

> It is fair to say you stole our multi-billion dollar intellectual output in that scenario.

No, we didn't. We simply took the model as-is.

The Chinese models are the result of just as many papers, GitHub repos, etc… AND the synthesized results of those.

> The Chinese models are the result of just as many papers, GitHub repos, etc… AND the synthesized results of those.

Right, but they aren't the ones whining that other people are getting "the synthesised results" for free.

I'm pretty sure the Chinese cloning isn't being arrived at for free regardless. Fable is quite expensive for example.

If Anthropic has a real problem with API use, they can always raise the price.

And you can download the Chinese model weights and run them yourself - admittedly not too practical for Kimi K3 unless you're a big corporation, but eminently doable for others. The hardware to run Deepseek r4 uncompressed is about about $30k no=ew, well within the power of a small company or financially secure individual. Compressing and/or getting creative with hardware could bring that down quite a bit.

The difference is that the Chinese are sharing the models with everyone.

Yeah real convenient that the line would be exactly after pilfering all of human knowledge work, but before American model providers.

... and they're at least releasing the results.

> comes from huge amounts of innovation

Thousands of years of human innovation taken without any permission.

Everyone should steal everything not nailed from other AI companies. Then steal everything nailed and take the nails too. At least this way a tiniest bit might return back to society.

a synthesized result that comes from huge amounts of innovation and computation

The published algorithms like the transformer architecture are not patentable. You spent a lot of money on compute and China used the uncopyright-able output to steer its own training models? Too bad. I feel especially unsympathetic to OpenAI, who went from being a presenting itself as a benevolent nonprofit to a very-much-for-private private entity over night.

You could say that both of them stole, but different stuff.

Don’t forget that the Chinese models are also built on top of huge amounts of “stolen data” as well, beyond the distilled. So it’s basically all of the above. However, there’s no mechanism for the NYT or an author or anyone in the US to sue the Chinese companies that took their work.

and if you're an author that lives outside the United States?

You can't steal intellectual property, only infringe on the copyright holder.

The Chinese models poked and prodded better models for training data and to avoid having to pay humans to RLHF themselves. You can call it infringement, theft, whatever, but it’s quite obviously “not ethical” to me.

Looks like these humans' jobs got replaced by AI.

Sorry, no leg to stand on and relatively speaking no sympathy for the model makers or the data-center builders…

I assume you don’t use any LLMs then

I use them and I'm fine with them training on our information and knowledge. It's what it's for. No one takes it or steals knowledge; they just copy it.

But because of that, I'm also ok with the Chinese doing it. The worst they might be guilty of is breaking a terms of service.

The only incoherent position is that it's good for one and not the other. You can consistently think it's bad in both cases, or good in both cases.

bet

The American LLMs have been equally distilled from Chinese ones. Not least because the people whose creativity in collecting training data barely extends to pirating Annas Archive probably lack in great Chinese datasets.

Try it yourself: https://imgur.com/ZfxYmaq

nice, you got claude to say "I'm deepseek" when queried/prompted in Chinese, that's great!

你是谁? -> 我是 DeepSeek 由深度求索公司...

You can't be blind to training costs. And you can't be blind to Meta dabbling in the openish strategy (Llama) before the Chinese labs did.

Whats even funnier is the attempt to restrict the hardware capabilities of Chinese models inevitably helped them (Because we know they're just as smart, if not smarter, than the staff in America) create smaller and leaner but just as capable models. That's why we now have upper-consumer models fitting on 24GB that can build, manage medium sized git repos. I've yet to find a git repo I can't throw at the Qwen3.6 35B and get it built and running.

So it's an endless amusement watching american capitalism do it's bloated oversized dance then get trounced by smaller, leaner activity. It's a pretty broad metaphor that is clearly poking at every american seam/.

Human ingenuity thrives on constraints.

Human sloth thrives on no constraints.

The compute constraints never mattered. If China had more compute they'd still end up winning because they have more people and a culture more inclined to math and science. Even if you find all this amusing, there's no own goal here. Not a policy one anyway.

It’s poetic.

[deleted]

It doesn't make me happy to say it, but the American LLM companies were first. Capital in the rest of the world is way more conservative, and I can't imagine the mega-investments OpenAI and Anthropic managed to secure happening anywhere else without existing proof that "thing is profitable".

First-Mover Disadvantage - https://hbr.org/2001/10/first-mover-disadvantage - October 2001

> In business today, it’s universally assumed that speed is good—that the fleet thrive while the laggards struggle just to survive. This belief is perhaps most strongly expressed in the concept of first-mover advantage. The company that leads the way into a new market, the thinking goes, locks in a competitive advantage that ensures superior sales and profits over the long term. It’s a nice theory, with a long pedigree. Unfortunately, the facts don’t support it. We recently completed an extensive study of the results turned in by market pioneers and followers, in both consumer and industrial segments, and we found that over the long haul, early movers are considerably less profitable than later entrants. Although pioneers do enjoy sustained revenue advantages, they also suffer from persistently high costs, which eventually overwhelm the sales gains.

First mover advantage is theoretically only a short-term advantage. Long-term revenues come from entrepreneurship, and a first mover may or may not better insight into long-term market wants than later entrants.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

Phones are constrained by battery power and memory does not shrink as fast as CPU/GPU, so unless there's a battery breakthrough and/or memory breakthrough, you're not fitting 100Gb of RAM on your phone in 10 years.

Absolutely in a Mac Studio equivalent.

LLMs have emergent capabilities when they get smarter. So who knows how insanely big frontier models might be at that time, or what their capabilities may be.

Not just that memory shrinks slower, it has practically completely stalled. On chip cache seems stuck at 7nm and DRAM is stuck at 10nm. As transistors shrink, they hold less charge, creating weaker signals that are harder to read and prone to interference. Smaller nodes aren't a huge issue on CPU/GPU work load because they don't have to hold a static state.

I'm not saying we are at peak memory but future gains are going to come increasingly slower.

You surely meant capacitors, not transistors

I think this is all true, but that unlike with Moore's law and improved PC tooling and capabilities, we also have essentially existing biological evidence that there should be a way to create much better intelligent systems in terms of training, memory, and efficiency. With classic PC evolution we didn't even have that evidence but still could make a relatively strong inference (Moore's law). But here we basically have evidence that there can be something much improved and know that it's only going to take research and discovery to figure it out, not new hardware processes.

More efficient AI is possible in principle but when the next breakthrough arrives there's no guarantee that it can be implemented on current hardware architectures. Something fundamentally different may be needed, as different from current GPU/TPU chips as they are from regular CPU/FPU chips.

The existing biological evidence took billions of years of evolution to get to the state it's at. So it may not just take research and discovery but also enormous computing power.

Evolution is really slow as you say. It even stops moving at all without selecting pressure.

Yes, one might call these… evolutionary algorithms…

Exactly. Even putting aside “better”, brains show that orders of magnitude greater efficiency are possible, for training and operation.

Indeed. It feels like we're at the equivalent of what the original IBM PC offered in the personal computer revolution.

Just seeing how much has progressed as far as capability in the past 4 years as far as capability and efficiency, it's clear that there's so much more to learn and refine from.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

At current pace, we'll have open weight LLMs with frontier intelligence in 6-12 months. The constraint is RAM - both for the model and the context. It's likely that distillation and quantisation and TurboQuant will significantly reduce RAM requirements. I think we'll have Opus 4.8-like performance on 64GB of RAM in two years.

Of course, by then, frontier intelligence will be god-like.

> god-like

So would you say we are months away from full self-driving cars that can out-drive a human being in any situation?

Remember, these cars are running local LLMs, not frontier models. The issue with self-driving cars has been the edge cases. The 0.0001% of situations where the models did not have sufficient training data. This is compounded by the hardware limitations. Onboard RAM in a typical Tesla on the road is 16GB (+16GB for the backup computer). This has to run the existing onboard OS and other operations plus the LLM. These two factors combined means that the cars are currently incapable of negotiating the 0.01% cases, let alone the 0.0001% cases. And this is compounded by the fact that LLMs cannot currently update their weights in real-time, like humans. It takes months to train a new model. Special small models can be very tricky, especially around safety and mission critical applications like FSD.

All that said, current data shows that FSD is already better than human drivers on average. See the recent regulatory decisions by the Dutch and Danish road safety authorities. So we've already crossed the rubicon. All improvements now are icing on the cake. My prediction is that local LLMs will get much better, very fast. How that's operationalised with Tesla (or other) data is yet to be seen. They have at least three new ASCIs/SoCs in the roadmap for improved LLM efficiency and with a lot more RAM. Plus they just announced new technologies allowing the local LLMs to learn from driver intervention and behaviour. Some form of vectorised RAG, which could mitigate a lot of the limitations around real-time learning.

I am very optimistic for the future of self driving. I own a Tesla with FSD now, and it's incredible. It makes mistakes, but fewer than I do, and so far has saved my butt (and my wife's) several times from obstacles and emergencies we would not have seen. The car has undeniably made us safer.

[deleted]

> so far has saved my butt (and my wife's) several times from obstacles and emergencies we would not have seen.

Honestly, you need to reflect on your driving habits. FSD has only been usable for two or three years maybe? And you already encountered MULTIPLE situations requiring active safety intervention to save you during this time?

You cannot rely on the extra safety it provides. A driver with basic competence should be able to avoid most risks through anticipation before they happen.

All of the scenarios involve circumstances we could not control. For example, someone pushing into our lane on the motorway in the blind spot, requiring evasive manoeuvre to avoid. It's possible Danish drivers are careless or aggressive, but I think it's more likely you just don't realise how many close calls and evasive manoeuvres you need to perform on an annual basis with average driving kilometers and regular defensive driving.

> You cannot rely on the extra safety it provides. A driver with basic competence should be able to avoid most risks through anticipation before they happen.

It's too late. These safety features are saving lives all over the world, every day. As per the Danish and Dutch regulators, FSD is already safer than the average driver. I would prefer that all the drivers on the road are using this technology than manually controlling their cars. People make all kinds of mistakes which FSD does not.

The value of an LLM is the dynamic reasoning you get out of it and the cost to execute on that.

I see two forces working against this that proprietary models will always have over an open source model.

1. The biggest is content licensing. Content is quickly becoming gated by systems at the front of their load balancers, completely changing the social contract of the Internet. What used to be a quick google search for recent facts that lead me to places like reddit or twitter, is now completely walled off if you're not physically at your browser and using an IP address from a last-mile provider.

LLMs have pre-trained on the bulk of the information up to 2024/2025, but over time that will be more and more out of date.

Anthropic, OpenAI and Google will all have to pay for access to a lot of this content refresh going forward, and it does make a material difference in the output you get.

2. Liability is the other. A corporation can look at a contract for model access and see one that provides uptime guarentees, content infringement promises and model safety, and pick the contract that shields the corporation from the most liability. A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all. They will simply bill you for time spent on their hardware and make promises that they won't log or inspect corporate traffic.

> A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all.

Why couldn't they?

> Mainframes survive, but serving a much tinier portion of the market than they used to.

I would argue mainframes rebranded to "cloud" which is ubiquitous and more people interact with this computer than any other type of device... only difference is that it's a browser instead of a terminal

There is constant shifting between client and server computation. I think it is a stretch to call cloud servers “mainframes”. There are still old school mainframes, running JCL, and old school mainframe DB2 and COBOL. That ain’t cloud.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

People overestimate what can happen in a year and underestimate what can happen in 5.

I'm betting that increased model efficiency and hardware optimisations will get us there a lot sooner. Biggest hurdle would be the memory prices though, if those do not drop back down it might take 15.

Define "winning."

Open source is cheap, yet its operating systems are the least-popular. But their existence is critical to a healthy market.

It's not zero sum.

What you're describing is how things become commoditized, but many companies are excellent at ensuring they aren't seen as commodities

[deleted]

Training cost is actually not that high — it’s fixed and amortizable across the lifetime of the model. Inference is expensive, and open weights don’t solve that problem — in fact, they might even encourage it, since a high cost of entry means consumers will pay for inference directly from the labs anyway.

Unfortunately it seems likely the winner will be the cloud providers. If anyone can run inference on open models, then profit will flow to the vendors who can afford the capital to run them. That’s the CSPs.

(It’s basically the same business model as pharmaceutical R&D, but the major difference is that nobody has even talked about patenting the models like a pharmaceutical company patents each new drug. I’m surprised about that, tbh — why give all the leverage to the cloud platforms? They aren’t training frontier models…)

Winners are hardware companies, GPUs,XPUs, HBM, memory, connectivity. Even CSPs are just compute renters, they charge a margin to make sure their hardware purchases can be made back. But given there are more and more AI CSPs, traditional, and neocloud, pricing competition is inevitable, and given the huge expense of hardware, CSPs are squeeze by the hardware companies and users seeking to lower their own costs.

Inevitably the CSPs will make their own hardware, especially as we start to see specialized chips for specific models or generic inference. This is already happening with Google and TPUs.

It’s easier for the CSPs to move into hardware than it is for Nvidia to move into cloud hosting.

Although as a middle ground I’ve been quite happy with Nvidia Brev for on-demand GPU instances from a select marketplace of CSP offerings. It’s a well kept secret IMO — great product (from an acquisition iirc).

CSPs making own hardware still needs hardware companies, they reduce the Nvidia tax but still need the likes of TSMC, Broadcom, micron/sk hynix, Marvell, the truth semi-companies. CSPs will not have the patents, IPs and talent to replace any of them.

Also, not sure how well CSPs inference stack is compared with vllm + nvidia. A lot of open weight models uses MoE, making the inference stack more complex.

True, though they could always buy one. I’m surprised this hasn’t happened yet, maybe due to anticompetitive risk? Google bought Motorola long ago which seemed to work well for their mobile device offerings at least.

> free and low-end eventually wins

Apple, the world's second most valuable company, seems like a counterexample.

Apple is free and low end given the context of what was being produced and sold to businesses decades ago.

The history of Apple shows that it's not. Consider what their early computers did to the industry.

> in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

The way things are going with regards to RAM/storage prices, I highly doubt that anyone but the richest among us will be able to afford them.

The ram crisis is just a short term bump. The three stooges of memory don’t have more than four or five years tops. This is their last big payday.

> - PC office productivity software destroyed expensive professional products.

I agree with the lesson too. Just to be precise, wouldn't the current model war be more akin to open-source office suite versus MS office suite? If so, then the cheaper option didn't really win. That said, the open-source alternatives didn't really feel the same as MS Office, and it took them a long time to reach the feature parity (or did they ever?). In contrast, the open-weights models are getting close enough to the SOTA models, and users can easily switch from one to another without feeling any difference for mojority of the tasks.

No, I don’t think so. PC + MS Office killed Wang custom hardware/software, for example. LibreOffice is much later, and can’t displace MS Office due to network effects. Cheaper won, cheapest can’t because of those effects.

I'm not clear what you mean with "PC office productivity software" but it seems like Microsoft Office is the winner there. Isn't that professional?

It didn’t used to be.

Phones are already running models locally which can be used in the field for specific use cases. Maybe not for frontier coding just yet.

Also you don't need to be connected to the network to use a local AI in many instances. If all mobile apps were done with a local-first approach, then you could use a local AI to query your emails, lookup already visited pages, summarise recently received documents, and lots more. Lots of apps could use an inbox/outbox approach for receiving and sending updates instead of relying on the network at all times. And this pattern could be greatly leveraged by local agents.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models.

I love the idea of SaaS offering these at lower rates today integrated into what ever you do and be 100% private. But I think the key challenge to mass adoption is productizing them in a way which makes sense for people to pay money for. As a commodity a local model is useless unless combined with some capabilities important to me. A PC is inherently useful because of so many applications offered on it on it. How local LLMs would be useful as a product that is useful for mass market is not yet proven.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.

Not likely. The last 50 years had Moore’s law growth in compute. That’s over. Frontier models are roughly compressed all written text and a large part of images. Those don’t compress forever, and likely not a ton more than now.

Inference requires touching a significant of that per token.

All of these are up against fundamental limits, more or less.

Moore's law is over in the literal sense but silicon continues to advance relatively quickly.

This claim isn't really outlandish in any way. It's not hard to imagine:

- Future models being able to handle current frontier models' workflows with much higher efficiency.

- Future consumer devices like phones having 2-4x the RAM onboard along with GPU/NPU performance greatly increased in 10-15 years.

Moore's law for 10-15 years was more like 20-100x ram sizes, not 2-4x.

Performance, storage, etc is definitely getting better, but it's a different scale of improvement

10 to 15 years from now the scale of LLM efficiency improvements is, quite literally, unpredictable.

It could be that the company valuations crash tomorrow, and (almost) only performance gains achievable on hobbyist-level hardware come to fruition from there on out.

Or it could be that in the future, we have a custom "model FPGA" à la Taalas [0] in every home, and that it turns out we can still massively boost inference efficiency due to novel discoveries like TurboQuant [1] or a somehow-improved quantization method [2] again and again ten times over.

Point is, Moore's law in this context shouldn't be applied to just hardware spec sheets alone, but more the total number of "parameters potentially improving", IMO.

[0] https://chatjimmy.ai

[1] https://research.google/blog/turboquant-redefining-ai-effici...

[2] https://prismml.com/news/bonsai-27b

[deleted]

It's probably not so bad. (I am not an expert on anything.) The big blocker is probably a roughly single-order-of-magnitude decrease in the cost per MiB of VRAM. (That obviously goes out the window if the LLM frontier people find, in the nearish future, new ways to do more with more: to significantly push up the threshold of diminishing returns from more VRAM or other resources. But that doesn't seem to be waiting to happen.) That's not clearly unachievable, especially given that ASML has apparently already made significant efficiency gains recently while there's no shortage of demand to justify R&D right now. Many customers are also likely to increase their hardware budgets: the kind of organisation that used to pay big money for Sun workstations is likely to consider spending that kind of money again if it saves them several hundred dollars a month in LLM plans.

I need a !remind me 15 years.

And yet the money is all at the high end. Bill Gates is much richer than Linus Torvalds. Oracle created one of the richest people (until he squandered it all on bad AI datacenter bets and got his company currently rated as a junk investment). Dell probably makes more money than IBM, but not by a lot.

That would sound very reasonable except this is the same what people said first about Windows and then Android crushing Apple.

And yet it's Apple that controls the top of the market and has the best margins in the business.

This is the same position OpenAI and Anthropic have right now.

Could this market be different? Maybe. But the status quo could be preserved as well.

I don’t agree. In the mainframe/PC battle, MS and Apple were basically on the same side. Once the dinosaurs went extinct, different mammals fought for dominance.

Counterpoint: linux owns the server space.

Cheap PCs enabled the creation of Linux. Linux runs on the descendants of PCs, not mainframes.

> free and low-end eventually wins

Not in SaaS which is what LLMs are. You can get VMs for much cheaper than AWS, Microsoft, and Google offer them but large companies (and startups) are happy to pay a premium for the support, reputation, and reliability that they perceive those companies as offering. Same thing for some of the managed database providers who are effectively selling a very heavily marked up version of postgres.

> The high price, and social pushback, mean that the American companies producing these models are precarious

I doubt it. The models really aren't that expensive when you look at what they can do. Fable is probably at least as good as the average software engineer and costs $50/wk on the max plan vs a software engineer who would cost closer to $4000 a week. The real money is probably in selling to enterprise vs consumers (Google has best route to making money from consumers since they can do what they did with ads and search to LLM queries).

It seems unlikely to me that US companies will send important corporate data to models controlled by a Chinese company as well.

> Fable is probably at least as good as the average software engineer and costs $50/wk on the max plan vs a software engineer who would cost closer to $4000 a week.

That's because the max plans are _massively_ subsidized. At API pricing the kind of usage to replace the value of a SWE is going to be way, WAY more than $50/wk. Orders of magnitude more. And to remain a frontier model org that kind of pricing has to continue in perpetuity.

There will always be a space for perforce in a world of git.

Doesn’t mean perforce is worth trillions.

That's not really my argument. It's that companies seem happy to pay a premium for a large company to provide complicated software services to them even when there are cheaper competitors.

Yes, as I pointed out, mainframes are still a thing. But the vast bulk of the market has moved on.

But aren’t the frontier models heavily subsidized? That’s not sustainable.

This really depends what you mean. If you spend $5.00 on anthropic tokens does that cost Anthropic more than $5.00 to serve? No.

If you amortize all of their training and salaries over that $5.00 then yes.

If you only amortize the training costs of that specific model then again we're back to no.

LLMs are not SaaS. Some LLM are delivered as SaaS. Anyone who has a machine big enough to run a frontier model can launch their own LLM SaaS tomorrow and it would be functionality indistinguishable from any other which is running a similar model...

Also, big companies can choose to run their own models on their own hardware and get better security and privacy as the data doesn't need to leave their own premises.

> Also, big companies can choose to run their own models on their own hardware and get better security and privacy as the data doesn't need to leave their own premises.

Yes, and then they would be reinventing the company owned data center that most big companies have just spent over a decade moving away from. I don't think companies will do that when there are multiple vendors competing to provide that service at what are quite reasonable prices when you consider what paying a human for similar output would cost.

How will open-source and open-weight models continue to thrive after financial incentives die off? Surely open models will suffer from outdated knowledge cutoffs if noone will pay for model training?

Nvidia will pay for training models that they can give away to consumers who will then buy their GPUs.

Pretty much. Nvidia has always been pretty decent at giving away software that is deeply dependent on their hardware.

A counterpoint would be all are chip fabs are in Taiwan right now due to huge investment. And there are lower end chip fabs around the world, but they have not cracked the major market.

What about AWS, Azure, cloud computing?

if you look at how GPU memory grew in the last 15 years, it's about 10x. Sadly, 10x from today doesn't get us to a typical frontier model size of today which is a quickly moving target. some other advancement needs to happen to get us another 10x both in memory/compute requirements, and also power requirements.

Capability per GB and per watt has also been going up lot. This will continue in the future as well (not necessary as the same rate as last years). But enough that I think Opus 4.8 level is reachable on consumer PCs within 10 years from its release. Say at the price point of 2000 USD in 2025 dollars.

I think you are right. One small exception I can think of is Microsoft Office still crushes Libre Office.

Apple seems a counter example, no?

The biggest exception is cloud. Big cloud carries an insane markup (bandwidth is like 10000X!) and everyone runs on it.

The strategy there is false openness where deployment complexity is the real proprietary moat. Sure Linux, Docker, Kubernetes, Postgres, and all the other standard tools in the box are open source and free, but they're also arcane and complex to run and hard to make fault tolerant. So you're lured in by "open" and then locked in via a kind of "death by a thousand cuts" complexity moat.

(Personally I hold the view that complexity and arcane-ness beyond a certain point is indistinguishable from closed in practice. Open source that's really complex and hard to run is not open in any meaningful sense.)

AI may not admit that kind of moat though, because AI is very good at slicing through that kind of thing. You can prompt a model to make itself compatible with another model or to change code to make it compatible. There's no moat because the moat bridges itself.

I think the whole idea of people running models is wrong. Models will run models, training will become distributed, and people will ask ModelNet to do whatever.

Computing tends to oscillate between centralised and decentralised models. It also oscillates between batch and timesharing.

Currently training is batched and centralised, access is timeshared and centralised.

But eventually a previous generation of computing turns into transparent networked infrastructure, and then you get another layer of new kinds of applications on top of it.

That's what happened with the Internet, and it will happen again with AI.

If Linux doesn't qualify as open, then what does?

At this point it's just a matter of having enough ram in your consumer computer.

Until we reach a terabyte of ram at affordable prices imho this isn't going to happen.

I agree, although I think it will be a lot sooner than 10-15 years. I'm running local AI right now and it's definitely not production grade yet, but it's surprisingly good. Speculative prediction that I probably shouldn't make: when the bubble pops, depending on when it pops, RAM prices might drop a lot. I could foresee these companies having produced a lot of RAM that suddenly doesn't have a buyer. (I know high bandwidth memory is different, but I imagine there are companies that will want to take advantage of that)

This is not really true considering Apple and nvidia are two most successful hardware companies, and they are notoriously closed. Not to mention microsoft, oracle they are all pretty closed.

You forgot smartphones, where low-cost did not win out. It led to low margins for the Chinese firms and eventually left them unable to invest properly in key markets. They may still hold marketshare, but in terms of profits, falls well short of Apple and Samsung.

I can see a lot of parallels here. Model performance doesn't matter if you can't make the system commercially sustainable.

> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

The parent comment cherry-picks evidence. There are plenty of counter-examples:

  * Office productivity suites
  * Search engines
  * Email services
  * Cloud services
  * Accounting software
etc. If the LLM market ends up like search engines, one company will dominate.

The obvious counterpoint: Apple captured the majority of the US smartphone market.

And no more than about 1/3 anywhere else

Libreoffice

> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

Is that actually true? There are very large markets that make a lot of money from paid software. And I would honestly prefer actually paying for software rather than constantly dealing with "not a bug" or "PRs are welcome".

Except cloud services won over local-first apps.

> Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) doing

I'm not even sure in 10-15 years whether we're still going to have consumer PCs, or PCs at all.

I have similar thoughts.

To those who feel on the contrary, I would genuinely like to understand why average consumer won't be priced out of hardware? The silicon industry is already quite centralised. Everywhere we already see the concept of ownership disappearing.

It's quite difficult for me to visualise a non-dystopian future where our PCs are just mere screens and every compute happens on a remote cloud, owned by some corporation, charging you subscription fees to even add and multiply numbers.

I would be the happiest if this (perhaps the most) pessimistic scenario doesn't pan out, but I can't deny that it feels like that's where we are heading.

We are heading for Evil Startrek.

It's quite difficult for me to visualise a non-dystopian future where our PCs are just mere screens and every compute happens on a remote cloud, owned by some corporation, charging you subscription fees to even add and multiply numbers.

I'm actually kind of surprised that hasn't happened by now even ignoring AI. Governments and marketers would love to be able to spy on literally everything you do, the copyright cartels would finally achieve their fantasy of full control over all hardware, and there really are benefits that it could offer to users (zero-effort backups, transparent access from anywhere, cost savings from dynamically switching from a single core for emails to many cores and a fast GPU for gaming).

we have already gone through multiple cycles of remote and edge compute (mainframe to pc, pc to cloud, cloud to phone). as the software/hardware landscape changes, the economics of what can be run on what device will change accordingly. I very much doubt that it will permanently go one way.

That’s not what this is. It’s industrial dumping applied to software. China has successfully applied this strategy to become the manufacturing workshop of the world.

If china is subsidizing training they diminish their off-shore competitors expectations of a viable return on investment. It’s trade-war behavior.

in 15 years we might all be fighting terminators

While I generally agree there's some nuance here and that is that there really are few new ideas. Old ideas just get recycled.

For example, mainframes and minicomputer. Yes they were displaced by PCs. But what is cloud computing if not mainframes 2.0?

I do agree that in the next 2-3 years we're going to see real growth in local LLMs as the hardware becomes more accessible. It won't even necessarily be cheaper because data centers can run 24/7 and have cheaper cooling and electricity. It'll be done for privacy because your prompts and responses are themselves a commodity to AI companies and they live under a legal grey cloud. For example, does AI usage break attorney-client privilege? There are lots of opinions on this but it hasn't been tested in court.

Windows - pfff, Microsoft charged companies like Dell for an install on computers they didn't even install it on !!!

paranoid me feels like the artificial gpu/ram/ssd shortages are a plot by the VCs to forcefully reclaim central control via new age mainframes

one certainly cannot buy a PC for cheap anymore

instead of 15 years I think it'll be more like 1.5 years.

I wouldn't be surprised if apple were shipping 512 GB unified RAM macbooks before 2030 and that would be standard issue for folks to use local LLMs for their daily work

With Apple’s recent history engineering and designing around companies that hinder their progress, I don’t think memory is going to be any different.

I also think the rest of the tech industry that can isn’t gonna be stalled for too long. This windfall will be the last for those three stooges of memory.

> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

Except, uhm, for ..you know, that one company that hit a trillion cap

But you're right: Just like how million dollar computers with 1 bit of RAM performing 1 operation a second and taking up a colossal cave were replaced by $1 laptops with a zillion zekabytes running at a trillion hertz (exact values may vary),

the sprawling data centers of today with a quadrillion GPUs powered by black holes will get replaced by breakthroughs in hardware and most importantly, algorithms:

The human brain is proof right here that intelligence doesn't require dinosaur-sized hardware or eat half the sun every second.

I actually wonder if we're seeing the limits of discrete binary logic: Maybe it's high time to give analog ternary and all that funky jazz an honest try :)

umm do Mac vs pc and Iphone vs android

Not sure what your point is. I’m commenting on a moment in time which is like an earlier moment in time. A different moment in time will have different analogs.

> The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.

Mac has won and it is not free.

Iphone has won and it is not low end.

It is low end comparing to mainframes

[dead]