I am not trying to defend Mozilla doing this, and I don't support sending data to cloud based services like this in a way that users won't understand.

But I also think that the state of the art in small LLM and user device capabilities aren't there yet to put a "good enough to be actually useful" local-only LLM as a prepackaged thing in a mass market distributed browser.

You don't want a browser that takes 10GB of extra RAM (on top of the memory hog that is having just 3 or 4 complex tabs open on its own already) and pegs your CPU at 99% usage for minutes at a time. And not in an era when mass market consumer laptops are still commonly 8GB or 16GB of total system RAM. Many of those with integrated-into-CPU onboard graphics (eg: not a gaming laptop with a discrete GPU on the PCI-E bus).

It'll be a catastrophe for laptop battery life, among other resource use problems. And an LLM that fits in under 8 to 10GB of RAM for CPU-only inference is not going to be nearly as capable as an off-device inference system.

I wish they had just done this with a very clear up front opt in (not enabled by default) thing that explains what Mistral is, that it's not some big American cloud company but a relatively small startup in France, and that your prompts/LLM interactions will go to their servers. And some documentation on how it will be handled/stored in a supposedly trustworthy manner.

Well, it depends on the task, doesn't it? "running shoes I looked at last week" / "Here's what I found in your browsing history:" doesn't need a 119 billion parameter frontier model; it's a RAG problem for the 0.6 B embedding models. That's an example Mozilla offers. Presumably to explain to their users why it's essential they hand over their last week's browsing history for this convenience (but it isn't! Hardly for that!)

I feel it's wrong to tell users that it's important and normal to relinquish all control of their—extremely personal—life history, in bulk, in plaintext, to strangers.

I agree wholeheartedly that remote server inference is super useful, and that local inference falls far short on many tasks. (I have no objection at all to Mozilla providing a cloud inference feature).

What I don't buy is that we must ask users to redraw their personal boundaries so that their most intimate life details, and remote frontier-model inference, overlap. They do not need to overlap.

You can accomplish a lot with private local inference with the smallest of models; and you can accomplish a lot on remote servers which aren't privy to everything. If some convenience is lost by not combining the two, well, so be it. I'm sure most people would agree, if all of this was laid out plainly.

You’re taking their shoe example way too literally. It’s just showing they’re catering to the average user, not an engineer. And tons of average Joe users now have very high expectations of how intelligent a LLM is because they’ve interacted with ChatGPT and the like. If there’s a super tiny model that is too dumb to do anything beyond shoe searching and they have to switch to Gemini and google AI search for anything more complex, this whole launch would be an immediate failure.

I'm of the opinion that local inference should be done to the greatest extent that is realistically possible, at the earlier time that the hardware/average user platform is capable of doing so. I personally spend a fair bit on kWh extra in my home electrical bill monthly for having a good sized chunk of local inference ability in my house, but that's not a common thing yet.

If mozilla is doing things to send users down the path of doing this externally, they need to be much more upfront and transparent with the users about where their data is going, and not bury it in some terms/conditions that only nerds will hunt for.

> "running shoes I looked at last week" / "Here's what I found in your browsing history:"

Does that need an LLM at all?

Not really, but it's probably easier to make it on top of LLM than to make specially-purposed tool for it, if we talking in terms of time-to-market effort.

SQLite FTS could have done this a decade ago. We've had good local search capabilities for two decades and they either been underused or abandoned (e.g. Google Desktop, Yahoo! Search). This may be a reasonable projection of where AI is headed. You can do a lot locally but there is too much incentive to centralize around cloud infrastructure, then the privacy concerns make that prohibitive and we end up with what feels like a false choice of cloud or bust.

Maybe it's LLM hype that will bring more powerful capabilities to the desktop?

I think classification and clusterisation is more important than search per se. "Where did I see that article about that weird psychological effect where people remember things they haven't seen" is not resolved by direct search, but can be - with some work - helped by language models and NLP. One doesn't need full blown frontier LLM for that, but bare word indexing would probably not do either.

Firefox is already using small local models for a few things, like tab grouping and link previews. I assume that this is for cases where a reasonably sized model isn't adequate.

> Firefox is already using small local models for a few things, like tab grouping and link previews

Which I immediately disabled alongside their "improved" Nova design, cause it doesn't improve productivity. These are all distractions from web browsing. If they want to improve anything, they should work on their abysmal spell checker, but that's not shiny enough.

Let's say the goal is doing something for work, so accuracy is important, and you prefer for it to be a nice reading experience, like a good translation.

As long as cloud model are somewhat better in these things, it's good that users would have the option to use cloud models.

Dont use this for work unless it has been verified by your corporate security and legal teams. Sharing business data with a company with which you dont have a enterprise agreement can land you in serious trouble.

I'll ding Mozilla directly then - Since your "we can't do this yet" approach is actually the literal thing Chrome is shipping...

https://developer.chrome.com/docs/ai/built-in/overview?_gl=1...

But the built in AI model in Chrome is 4GB of RAM, 4GB of disk space, and is still only used in a couple places

> that it's not some big American cloud company but a relatively small startup in France

How it's any better? Small companies can be bought by big companies. French government is as capable of trampling over their citizen's privacy the moment they feel they need it as US government is. And putting a squeeze on a small startup is way easier than on a major cloud company (not that either is particularly hard). Also, OpenAI used to be an idealistic non-profit one day too, then it started to smell trillions and all that went of of the window.

> is not going to be nearly as capable as an off-device inference system.

I rarely need PhD-level research into my browsing history. I'm not going to solve millennium problems on my bookmarks. The tasks that I will realistically need are well within capacity of most very basic local models. Maybe they'd be a bit slower, who cares.

Let's not pretend that EU data laws are not far more stringent and actually enforced than whatever the US has. Like sure, the government with a jury decision may acquire some person's data. But it's not getting sold to whichever-random-company-offered-the-highest.

Also, some shady TOS is also not enough to override the law in the EU, you have much more protection as a user here.

> But it's not getting sold to whichever-random-company-offered-the-highest.

Nothing in GDPR prevents transfer or selling of personal data. True, there are hoops to be jumped to do that, both in terms of consent and documentation, but let's not pretend a lot of people would read the welcome banner and refuse to interact with a service because its legalese says "we'll sell you data and if you don't want us to, go away". Yes, selling the data becomes more expensive because you need to hire the lawyers to produce those welcome pages and regulatory compliance paperwork, but large corps have enough money for lawyers.

> Also, some shady TOS is also not enough to override the law in the EU

It doesn't need to override anything, as the law does not prohibit data collection or transfer. It only describes the hoops that need to be jumped to achieve it.

It certainly does. You need to weigh the users privacy with your interests and can’t just put them aside. Noncompliance leads to quite large fines. You can easily find examples of fines given.

The DPA’s are short on capacity, sure, and large corps will try to fight the fines in court. But to describe it as “just hoops whilst allowing everything” is unfair.

> You need to weigh the users privacy with your interests and can’t just put them aside.

Who said anything about putting them aside? There would be a process and highly paid lawyers confirming that selling user data is only for the user's ultimate benefit, as it allows to provide awesome services to the users, and it all will be outlined in a 100-page privacy policy which you will be sure to read on each site you use, wouldn't you?

> But to describe it as “just hoops whilst allowing everything” is unfair.

Can you quote me the place where GDPR prohibits this? Not says something like "weigh user privacy" and "take adequate measures" and so on - which can always be resolved as "we weighed carefully and we took measures and we decided selling the data was awesome and users love it" - but explicitly and unambiguously prohibits the practice? If you don't find it - that description is exactly what it is.

>How it's any better? Small companies can be bought by big companies. French government is as capable of trampling over their citizen's privacy the moment they feel they need it as US government is.

It's better for Europeans because the company in question is European and so will not export their data to New Jersey. And because any success it has will presumably better benefit France and Europe.

> It's better for Europeans because the company in question is European and so will not export their data to New Jersey

There aren't many government institutions in New Jersey that would take a lot of interest in anything Frenchmen are doing. There are a lot of institutions in France that would be interested in anything Frenchmen are doing, if it contradicts what the French government wants to be happening. The main threat to citizen's privacy always comes from the government closest to them, the government overseas has its own citizens to worry about and pays much less attention to foreign citizens on foreign land. There could be exceptions, true, but as a rule, if you look into New Jersey, you'd sooner find mass surveillance of American citizens than mass surveillance of the French.

Same goes for commercial interests. If I want to run targeted ads in New Jersey, I want to have profiles of New Jersey people, not French people. So buying data from France in New Jersey would not be a routine occurrence, but buying local data would be.

> How it's any better? Small companies can be bought by big companies.

This. I've come to view a startup as a company without a business model, doing everything to get acquired by a mega corp that will finally squeeze the juice out of the userbase

(The lack of a viable business model applies to some mega corps too)

>But I also think that the state of the art in small LLM and user device capabilities aren't there yet to put a "good enough to be actually useful" local-only LLM

so then don't add it in as highly advertised feature until it is. doing things right and living up to your core values is a lot to expect from businesses these days but, at minimum, a non-profit foundation should be able to live up to these goals, yes?

But apple said their AI only falls back to the private cloud when it has to? Just kidding, near every request needs to fallback, because a phone can't actually run a real LLM, no matter how many "neural cores" it has..

Agree, advocating for Firefox developing features only for rich hobbyists(people who can afford large RAM and GPU) is absurd.

>advocating for Firefox developing features only for rich hobbyists

?:

>marketing pages aren't candid enough to clearly explain

The only people having an expectation of translations being done locally is exactly the nerds that keep whining that it's not using a local model. Every single normal person, when presented with a "translate" button either know it's going online, or don't care about it.

Begging the purists to run away from Firefox at this point so they can stop wasting everyone's time. Your demands for examplarity and whining about money not going ONLY to firefox and jerking yourselves on Servo was not enough, now you want to restrict the browser to owners of an RTX5080 if they want to use it?

> only people having an expectation of translations being done locally is exactly the nerds that keep whining that it's not using a local model

This is an untested assumption in Silicon Valley. I suspect Apple is going to eat a lot of folks’ lunches.

You realize that people that don't care won't even know what Firefox is? Being a niche player and scaring away niche people is the most stupid strategy possible.

Being a small player and throwing away any hope of expanding your user base by building for people who will complain no matter what you do is, well, something!

neglecting existing customers in an attempt to gain larger market share is certainly a play

Ah yes, the customers of Mozilla, who are certainly paying users and not just people whining and saying BACK IN MY DAYS IN THE LAST MILLENIUM IT WAS DIFFERENT.

Get over it gramps, things are changing. Mozilla spent ten years trying to mostly appeal to these people by staying very conservative, and it got them nothing but abuse and harassment from people not realising that staying stuck in the past is a death sentence for Mozilla.

You want to save Mozilla, burn down Google first. The ecosystem in which they are evolving isn't driven by them.

I solemnly proclaim I would pay reasonable money for a browser (Mozilla or otherwise) if their maker would truly try to make the best browser possible and succeed at least partially at it, instead of serving a thousand of agendas none of which is making the best browser possible. The sole reason why I am not Mozilla's customer is exactly that - because they do not see me as their potential customer and don't care about my needs. Of course, they are free to do that but the result is people like me aren't giving them their money. So they have to go to Google.

like I said, neglecting existing customers in an attempt to gain larger market share is certainly a play

the rest of your post seems to hinge on some judgement you think I made about that

> You want to save Mozilla, burn down Google first.

Uhhh... 85% of mozilla's revenue comes from Google via the search deal. In a theoretical scenario where a genie waved a magic wand tomorrow and google went "poof", disappearing, mozilla would be in a dire emergency for lack of revenue.

It would. But Mozilla has zero long term survival chance in the current ecosystem. The economics are just not there, there cannot be more than one big browser in the way this system works.

[flagged]

Whoa - you can't attack other users like this on HN, no matter how wrong someone is or you feel they are.

We ban accounts that post like this, as you surely know. We don't want to ban you, so please don't post like this!

https://news.ycombinator.com/newsguidelines.html

I do not fundamentally disagree with you but it's also an extremely well known phenomenon that average non tech users, in the aggregate of millions of people, will click almost any "yes/I agree/continue/Next" step on a software installer or new user sign up workflow for anything, without reading the ToS. People blithly sign up for all sorts of cloud based things and services without understanding their full ramifications all the time.

I also wonder at the specific level of aggression and the tone of your comment which does not seem to be an appropriate response to that person's specific comment.

People such as you are describing and rightfully criticizing are knowingly taking advantage of that. Indeed it's how a lot of malware gets installed too.

The person you're responding to is pointing out that a lot of people at the surface level do only appear to care about the results. They put something into google translate, it works, they gets results they are pleased with, they don't put a lot of thought into the fact that the data is going to an external service. That's not an inaccurate description of how a lot of people use their computers these days. Look at how many people signed up for ChatGPT accounts and put the chatgpt app on their phones and talk to it all the time. That's the level of critical thinking a lot of non tech users have about their personal data.

The fact that people will click yes/agree/OK on almost anything is how Windows computers got Bonzi Buddy installed on them back in the day, and now it's continued into the cloud-everything era.

https://geekhack.org/index.php?action=dlattach;topic=21140.0...

> it's also an extremely well known phenomenon that average non tech users, in the aggregate of millions of people, will click almost any "yes/I agree/continue/Next" step on a software installer or new user sign up workflow for anything, without reading the ToS. People blithly sign up for all sorts of cloud based things and services without understanding their full ramifications all the time.

I would put a lot of that down to learned helplessness.

A lot of non-tech people tell me that "they already know" and that it is impossible to avoid. I have been told I am naive to think its possible to keep data private.

It's almost as if outside the tech bubbles most people don't understand the tech or implications. HN posters who live and breath in this space are so smug to pretend this esoteric knowledge about networking identifiers or fingerprinting is simple and that anyone can get it. These are the same people who will throw their hands up on any other esoteric topic and give up on it equally. For example, tech people so easily decry how "stupid" the law and legal jurisprudence is (because obviously if they don't understand it, it must be bullshit) but see zero irony in being smug pretentious assholes to regular people who didn't happen to major in computer science.

We shouldn't all have to be PhDs in law and jurisprudence to trust a legal system to look after our interests just like we shouldn't all have to be computer scientists to trust that technical stewards are not exploiting people's lack of knowledge.

Sick of this mentality. Oh these people are sheep idiots who don't understand the EULA euphemisms of "improved user experience " means adtech and behavioural tracking.

I think youre right. The least techy people I know where the first to ditch Mozilla when they started to see ads in Firefox.

Can you please not go on enraged tirades on HN? You've posted several fulminatory comments in this thread alone:

https://news.ycombinator.com/item?id=49729021

https://news.ycombinator.com/item?id=49728539

https://news.ycombinator.com/item?id=49728246

https://news.ycombinator.com/item?id=49728175

and at least one more in another thread today. This is not what HN is for and it destroys what it is for. You may not owe Mozilla better, but you owe HN better if you want to participate here. This is only a place where people want to participate becuase enough people make the effort to raise the standards rather than drag them down. https://news.ycombinator.com/newsguidelines.html

I love this site

PREACH

That's a lot of words and assumptions, when a single check of my posting history would show that I despise the HN bros as well. Jumping on a tangent about libertarianism when you could have simply called them retarded and saved a lot of time.

Anyways, no, I'm talking about the average person, the public worker, the person that thinks the internet is the funnily named Safari app, the elderly: they give zero fucks about it going to some service online. They used to search for Google translate before, whether or not it goes on someone's server, they do. not. care. You're not going to win them over with "it runs on your device". Their device is a crappy laptop that barely runs excel, and if they can offload computing, they will.

They're advocating for a CHOICE and for the difference to be explained.

Have you tried Ling-3.0-tiny? It runs fine on CPU-- on a 14700KF gets 40tg/s and 250pp/s and on a ordinary gpu (RTX 4070) does over 200tg/s with no MTP and 7185pp/s. (my figures are Q8, though presumably a good Q4 would be faster)

It's certainly not as capable as something that needs a high memory gpu for quick performance, but I was quite impressed with it for what it is.

(and fwiw, I had it translate your last paragraph to German, then used google translate back to english: "I wish they had handled this clearly and transparently via an opt-in mechanism—not enabled by default—that explains what Mistral is (not a major American cloud company, but a relatively small French startup) and that your prompts and LLM activities are sent to their servers. I also wish there were documentation explaining how the data is handled and stored in a way that inspires trust.").

How much RAM does it take up in total? I'll have to give that a try on one of my test systems. Looking at a somewhat randomly chose GGUF quantization of it, looks like just under 5GB on disk in Q4, so RAM usage somewhere around 5-6GB?

https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF

That sounds about right for Q4.

it's an extremely sparse MOE, so there is some odds of acceptable performance using a smaller in-memory cache and the rest on flash. ... I don't have a setup to test that right now.

(Of course, if translation is all you want much smaller models will work. Ling-tiny can do summarization, dom manipulation, scripting, etc. too).

Andreessen Horowitz led Mistral's €385M Series A in December 2023.

Is that like 2-4 RTX9000 cards price rn?

> small LLM

Read this again, slowly.

> You don't want a browser that takes 10GB of extra RAM (on top of the memory hog that is having just 3 or 4 complex tabs open on its own already) and pegs your CPU at 99% usage for minutes at a time.

But I do, especially when the choice is either having it locally, remotely, or not at all. I also indeed do nit my CPU used for such, but GPU. I could even run a 8 GB model on a remote (but still local network, on-prem) NPU.

There's one caveat though: if you are gaming and browsing.

> But I also think that the state of the art in small LLM and user device capabilities aren't there yet to put a "good enough to be actually useful" local-only LLM as a prepackaged thing in a mass market distributed browser.

Then just allow it to be enabled on high end devices? But it must be local only. As hardware advances and people upgrade, more people will be able to turn on the feature.

[deleted]

If the thing you wan to do is not possible without totally compromising your ideals, maybe don't do it ?

> I am not trying to defend Mozilla doing this

Well, you kind of are though.

> I wish they had just done this with a very clear up front opt in

If a local model is not realistic, then this should not have even been an in-your-face opt-in, but at most some add-on.

Of course, their telemetry isn't even opt-out, so even the opt-out for the Mistral thing is kind of disingenuous on their part, since they get a bunch of information from us in other ways.

(sigh) Ah, Mozilla has gone down such a dark path over the years. Too bad.

I'm worried that the middle could fall out of the computing market across the board. If you can afford to keep up with the upgrade treadmill, you'll get private, local inference capabilities. If you can't afford to stay on the treadmill, you'll be stuck with whatever cloudshit malware Silicon Valley wants to foist on you.

I acknowledge that this is already the case, to some extent. The cheapest laptops at Best Buy are crammed with the most preinstalled malware. That's been the case for, what, 25 years? But you've always been able to wipe that cheap laptop and make it into a much more capable, trustworthy machine.

Well, assuming LLMs do become a pervasive part of the computing experience, what happens to the cheap laptops? Do all computers get more expensive to accommodate local inference? Does the rift between the everyday user's experience and the savvy user's experience grow even wider than it already is? Neither outcome seems good for the average joe who just needs to check his email.

Remember how in like 1999/2000 Sun was trying to predict that everyone's computer would be some form of thin terminal in the future? Turns out they were very wrong on the part about it running on Sun server back-end infrastructure, but that same general purpose has now been accomplished through other methods where a lot of people do basically EVERYTHING inside a web browser tab to some external cloud service.

Now add the need for external inference because very few random consumers are going to buy a $3000 laptop when they can get the $600 laptop at Best Buy, and that trend further escalates.

This has always been the case, forever. You have to pay for a product or service. How you do so can be with cash or your data/body/vote/eyeballs/indirect discretionary purchases.

The amount of work that can be done funded by foundations and free work is nowhere close to what people want.

Oh give me a break. This argument that people's objections to advertising comes from some Pollyanna naivety over things being free is nonsense. Tell me the last time you have ever seen a company be upfront and explicitly offer a free tier where they tell you upfront exactly what they are harvesting about you and for what purpose (and no, "improving user experience" isn't being upfront) but also offer you a paid version where they explicitly promise not to do that.

"If the product is free then you are the product" is supposed to be a cautionary observation, not an axiomatic proscription.

Does anyone actually trust companies that offer paid services to not harvest their data? Whenever I see this trope I think to myself that whatever lack of regulation, oversight and enforcement lead to that being okay would equally allow for them to both take my money and harvest all my of data anyways.

This is akin to people who criticise socialised medicine or services and pejoratively characterize supporters as just wanting "free stuff", as if the concept of collective payment is naive or something

Every ounce of RAM and spare cycle should be used.

[dead]

[flagged]