"I just wished the collected data was public. "

That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.

I agree. The double standard is the problem. People have been imprisoned for IP theft, but when these companies commit IP theft on the grandest scale ever imaginable, they're rewarded with trillion dollar IPOs. Either IP isn't protected, or it is. Legislators need to pick a lane. Right now it appears that poor people go to prison, and rich people get rewarded.

IP does not protect the little guy. This is nothing new. Draw a picture and then people start putting it on t-shirts and posters without paying you? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Self publish a book and then people start uploading PDFs of it? Better hope your real passion is filing takedown requests instead of writing.

[deleted]

AI training is the clearest example yet that companies are allowed to get away with what is treated as a serious crime only when an individual does it. There are many more examples of this, of course, but this one seems to be the most stark and obvious.

Legislators have consistently picked a lane. Protect the rich and powerful.

Huey Freeman: "Kim Dotcom was pissed."

Other companies have no qualms about distilling the first. Let's hop on gear and get the market to deliver a distilled Fable that runs on a smartwatch. Sooner is better.

I don't think this trend of open sourcing LLM will continue for a simple reason: Money.

I think the difference is that the companies are dumping billions of dollars into transforming that data into something useful, so they would like a return on their profits. Opening up the models for free is not a good business model if you want to make money.

I think the difference is that people are investing significant amounts of time, effort, and money into transforming their work into something useful, so naturally they would like some return on that investment. Giving away that work for free is not a particularly good business model if those people expect to be compensated for the value they create.

There’s too many implicit assumptions in that sentence that run afoul of the conversation.

Just spending money doesn’t mean it’s legal, for example. Criminals expect RoI too.

Anthropic getting angry other AIs are trained on their AIs output is, to me, one of the stupidest things I’ve read in a while.

On the flip side, output from an LLM is not copyrighted.

That has not been decided. The only thing that's been decided is that the LLM itself does not have copyright on its output.