Fully open models really need to be a big part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation.

Other than open training data (currently legally impossible), all of this holds for basically every major Chinese-made model. They not only open the weights but publish detailed methodology papers alongside the models in arXiv and even open source the code.

Inference code, yes, but the specifics of their training process (as well as the training of the vast majority of all other open weights models) are still a complete blackbox, and I can't think of any Chinese model that made its training corpus public.

They don’t release all the code.

The training data would need to have a permissive license for this to be possible.

Or, we just need to get this over with and declare any digital data findable via the internet to just be public property of everyone. Everything becomes public, besides stuff you keep locally, and there is no difference anymore, it's all just data anyone can use for whatever. A 1 year grace period for everyone to pull stuff off they don't want to be a part of this bright new open era, then we just scrap everything related to intellectual property, copyright and similar stupid stuff, and slap UBI on top of all of it for good measure.

I'd prefer to stay within the [hacker ethics](https://www.ccc.de/en/hackerethics), and protect private data. For non-private/personal data, sure. But individual people need their privacy protected.

Me too, I'm hacker ethics all the way, which is why I'm saying anything network connected should really realize the "All information should be free." dream, and then private data should be far away from the internet, on computers/drives not even connected to the internet. The whole E2E encryption is a ticking time bomb people rely to keep their data safe from others, but nothing that you don't physically have close to you can be truly secret forever, and even then it'll be hard.

I don't really get what you're suggesting. You give a 1 year grace period for Metallica to pull all its music off the Internet, but then as soon as I host some of their MP3s on my Wordpress blog it's "public property of everyone" from that point forward?

I hate to invoke Poe's law but, I've now flip-flopped like six times over whether this could be serious.

I think it is serious. In which case, I gotta say, it really seems like you didn't spend much time thinking about this. "A 1 year grave period for everyone to pull stuff off they don't want to be a part of" - How does that work when the Internet is already full of unauthorized reproductions, most of which people aren't even aware of? Even ignoring practical considerations, when literally everyone is basically stuck using the Internet for everything, this seems a bit unfair to anyone who isn't onboard, akin to The Onion's Google Opt-out Village. But there are so many practical issues with this, it would be easier to list the number of problems this doesn't have. You accidentally leak something to the Internet and it becomes commons? What happens when other people leak things to the Internet? How about revenge porn?

Not minor stuff that can easily be papered over, this literally reintroduces the problem of needing to care about the provenance of data again, in a way that can't be automated, which makes the whole thing entirely moot. All just to make training data for AI models easier to distribute?

I'm all for intellectual property reform, maybe even fairly radical. But this just seems like it wasn't thought out.

If this was satire, well, I took the bait. Oddly convincing despite being hard to believe.

[deleted]

What you're slightly more realistically looking for here is for publicly available data to have a Fair Use exemption for certain uses, which is certainly something worth discussing.

I respectful disagree. I enjoy reading e.g. Asimov and well-executed journalism. And I completely respect the IP of those people who create these works.

I'm what way does it make sense to respect the intellectual property rights of a dead man?

If IP rights end at death, there are some significant perverse incentives for offing big-name artists and authors and whatnot.

The current "lifetime of the author + X years" rules in effect in the United States still carry the same perverse incentive, though the incentive diminishes rapidly as X gets larger; with the very large value of X in effect today the perverse incentive is so small as to be effectively non-existent, but it's still there in theory.

Personally, I'd prefer a fixed term. I know enough independent authors making a living from selling their books that I'm willing to allow the fixed term to be large, like 50 years from date of completion of the work. (With a good definition of "completion" so someone can't cheat by editing a couple lines per year to keep something copyrighted indefinitely). The simpler the rule is, the easier it is to understand, and the harder it is to cheat it. The more complicated you make a rule, the more loopholes get found.

There are significant perverse incentives for me shooting you with a gun and taking your money and running away too. It mostly doesn't happen.

That dead man took the risk of not earning much in his lifetime, to continue feeding his family even after his passing. You might as well ask why does someone acquire life insurance.

You could sidestep it by running non-permissibly licensed training data that you purchased through an LLM. Legal attitude so far seems to be that this is transformative as long as it's not 1:1. The question on whether or not the end result is copyrightable of course remains controversial and inconsistent, but that question is also fairly irrelevent. You don't get more libre than public domain.

That's a fair amount of computational and labor overhead mind you, as you'll need to verify and prune the quality of your mountain of synthetic data, but certainly possible.

Though this assumes the legal system is a rational actor playing by the set of rules it claims to. In fact, I highly suspect you could get very unlucky and get an unfavorable ruling against you, because you stepped on a big pile of money's toes in the process of doing this.

It can also be used to sidestep copyright like this forum, books and most websites even if the data was not purchased but is a website or book.

Are LLMs what we need to make all data public domain? This way it could be used for that purpose

Hear me out.

Decentralized unstoppable storage, combined with decentralized unstoppable training, sorta like SETI for AI training. The seed of this tech already exists with IPFS and others like it.

We know (some? all?) of the big labs have skirted copyright laws at one point or another. Truly open models would just build on what is publicly available.

[deleted]

Crypto bros took the idea with some blockchain shit and no one takes it seriously anymore so it died

If the LLM/AI ecosystem starts actually needing some Person-To-Person (or maybe Agent-To-Agent?) payment system because things actually get smart enough to be useful autonomously, they're gonna need some way to send money/currency around. Depending on how banks will react to this need, we might see another return of digital currencies from the current winter.

They used to say that cryptocurrency will be the dopamine layer of the first artificial intelligence

UAE's IFM / LLM360 MO is indeed "fully open source" LLMs: https://www.llm360.ai/reports/LLM360-Towards-Fully-Transpare...

Eventually we'll just construct 100% synthetic training data that can reliably reproduce pretrains and fine tunes.

The first broadly useful fully open source models will do this.

We already have open data / open code / open weights for some domain-specific cases, such as audio models trained on large open datasets, eg. Tacotron / LJSpeech from waaay back in the day, though that is certainly not SOTA anymore.

Distillation could possibly be considered an early case of this as raw AI outputs are themselves not copyrightable unless humans enrich, filter, or transform them. Granted, that does not handle the cases where the outputs are sufficiently similar to copyrighted original works.

Where does that synthetic data come from? Magically just started existing?

But how much of that synthetic data still ultimately derives from non-open sources? You'd still have to ask what a clean room implementation ultimately is, depending on how granular or aggressive a large publisher wanted to get about it.

That said, I don't necessarily disagree with you. Talkie[1] presents an interesting case for it being at least possible to do this entirely on public domain material.

But even that used Claude somewhere in the course of its training pipeline (it's listed as a contributor on their GitHub), so again, how granular you want to get with that is still a question.

[1] https://talkie-lm.com/chat

Why? Sure, I’d prefer it, too, but this is just another GNU/Linux vs. macOS situation: most of us would prefer the first, but actually get shit done on the latter.

Why do you make the worse choice and not use what you would prefer to use? You have been able to "get shit done" on Linux for nearly 30 years. Have the courage of your convictions.

Without open-source, there'd be no macOS.. So good thing, it exists.

And that's why companies shouldn't fear opening up, but having both is still a net benefit.

which is why everyone runs docker on mac, to get shit done.

we get shit done on the cloud with the former rather than the later

I personally find the analogy unconvincing, the UX dimension is completely different as I can use the same harness with any model; and the year of the linux desktop is coming soon (tm)

I believe Olmo from AllenAi is this

https://allenai.org/olmo

Open models can be used/changed for social manipulation too, by anyone, which scares a bunch of people, as opposed to the dark pattern manipulation from Big Ai/Tech

But that would be impossible due copyrights laws. If the law would apply Anthropic and OpenAI executives would be in jail

Money is the issue here, no one wants to fund it.

I'm sure anthropic didn't want to fund the extra "safety" guardrails they put into fable, but they were forced to, else they couldn't release it.

Sure there are all kinds of problems with that situation. But it still demonstrates that they can be coerced: play nice or don't play at all.