I just don't understand people saying "but a human learning from a book isn't illegal".
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
Because it’s a bad faith argument that presupposes integrating someone else’s work into your algorithm is equivalent to me reading a book.
Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.
It's the same bad faith argument as:
"It's perfectly OK for a police officer to observe a street corner, see crime happening, and go take action, therefore, building a complete, panopticon surveillance system that watches all street corners simultaneously and deploys police to take action, is perfectly OK, too, since that is exactly the same thing."
I haven't seen this articulation before, and I'm going to remember this one!
In return, I'll share my own reasoning:
Human Life is finite, and there is a real opportunity cost to learning (say) to copy Picasso's style -- you could've been doing something else in that same Time.
However, if you're training a model on the entire career output of dozens of artists, it only needs more electricity + GPUs to do this.
And, transposing this back to the Human realm, there's no way one person can put in that kind of wide effort to learn how to copy dozens of artists' styles in one lifetime.
I really like the way you put it too -- it's a bit more succinct.
The quantity of an action can have a qualitative difference. I wish I didn’t haven’t to explain that to so many people everytime a new technology is created.
Quantity has a quality all its own
I don't know if it's even OK for a cop to just randomly watch a random street corner without a specific reason to do so. If a cop was following a given person around all day every day without probable cause would this not be considered harassment? If we view the Flock Camera network as a single system is this not harassment?
Just to add context and not to comment on rightness, that was one of the early core purposes of policing--to ensure that a people in a particular public space (where either the people or the space are particularly vulnerable) are free from unwanted disturbance. Here is that exact concept shown in a painting from 172 years ago: https://en.wikipedia.org/wiki/The_Gleaners_(Jules_Breton).
Correct take according to law. A mechanical style royalty would be requisite
[dead]
"How do people not understand that some laws only make sense at a certain scale? "
Honestly, this is because its something that is basically never discussed or reasoned about. The closest I can think of is "personal use vs commercial use". But I'd love to see more about how to reason about how laws change at scale.
Laws do consider scale and we can see that in everyday law.
4 friends walk together, it's just normal. 400 "friends" walking can and will be treated differently.
Moving around with a couple of bills is treated differently from carrying huge bundles of cash.
Those laws or exceptions were probably added later as a reaction to abuse of existing laws.
The problem with AI/scraping is that we can't afford to be reactionary because it may be too late by the time we realize what has happened and change the law.
It will be too late because governments and judiciary have been largely about maintaining the status quo and minimizing disruption when it comes to big tech related cases even when they have been found guilty of wrongdoings. We already see the too big to fail vibes with AI.
Laws are typically implicitly considerate of scale. Very rarely are they explicitly considerate of scale.
To wit, sharing music digitally, Google scanning books, or Uber/Lyft providing unlicensed taxi services.
It's pretty rare that laws consider what should happen if it were suddenly possible to 10x or 100x preexisting throughput.
And specifically where laws balance multiple, often-opposed, stakeholders' interests, that change can drastically upset the previously negotiated compromise.
Which is why piracy at scale, before it's banned, tends to be a successful foundation for many businesses.
Murder, terrorism and genocide.
Simple possession of drugs vs possession with intent to distribute. In jurisdictions that make difference base on quantity (so, scale) alone.
I'm sure there's other examples too.
Theft was historically very scale-sensitive. In England you'd face execution for "grand larceny" if the goods in question were worth more than twelve pence. This was on the books from the 13th century until 1832 (!).
Jesus Christ on a pogo stick, execution for stealing 12 pence worth of stuff? That's crazy! Adjusting for inflation[1][2], that would be the same as nowadays being executed for stealing £5!
[1] According to https://www.bankofengland.co.uk/education/education-resource..., 12 pence (1 shilling, or 1/20 of a pound)) in the old £sd system is equal to 5 pence (or £0.05) in the later decimal system.
[2] According to https://www.bankofengland.co.uk/monetary-policy/inflation/in..., £0.05 (decimal) in 1832 would be equivalent to £5.02 in 2026 Aug.
When first introduced it was around the price of a sheep. The homicidal version of bracket creep [1].
[1] https://en.wikipedia.org/wiki/Bracket_creep
Damn, sheep were cheap back then.
Because copyright is a tradeoff. Society wants information to be free. But authors won’t publish if anyone can republish their work for free. Hence the distinction: The ideas you write are not protected - only your expression is.
In a sense, AI changes nothing. Society profits from having AI just as it profits from having people learning from others. In both cases, those who stand on the shoulders of others still make money for themselves. But the economy as a file is richer, too, because people can choose to buy something better now that wasn’t available before.
Allowing AI companies to get away with this because we're "better off"¹ is like eating our seeds. Sure, we built something cool with the massive corpus available, but how will it affect future decisions to develop or share creative work?
If nothing else, a sense of justice tells me that if somebody's work directly helps create a profitable tool, that person should share some of the profit. The size of the share can be negotiated, but the AI companies didn't even reach out before the law suits. And even then, only to major sources of content (some of whom don't have the copyright for their content, just a limited license for distribution on a website and all the nitty gritty involved in that).
1: In a sense, a world with these models has more capabilities and is therefore better. But this is the real world with real people, who are emotional and competitive. So let's see how it actually plays out.
Agreed. The most damning behavior was Meta.
Specifically, reaching out and discussing licensing material, then pirating because it was too expensive/slow to legally acquire it.
Heaven forbid Meta have to pay for something.
> authors won’t publish if anyone can republish their work for free
There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).
[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.
Again, incentives.
The people who publish are people who have reason to publish when they can be copied. Typically either they have already been paid, or they expect to gain market share by being free.
People who need renumeration to continue working will not publish.
Intellectual property rights, as much as I dislike the RIAA and MPAA, created a way for more players to enter the market, because it created a way for their needs to be met.
>Society profits from having AI
Not all of society at all. Let's discuss this when AI is better integrated and a large fraction of people are laid off in 5 years.
I really don’t understand why anyone would think they are entitled to any part of a derivative of a published piece. Why publish if not to help advance humanity, just like all who came before you and contributed to your success/ability? This is quite literally the bedrock of the progress of human civilization. As long as they aren’t straight up reproducing a direct copy of the original work.
If you want to keep it to yourself so only you can benefit from it, then keep it private. Otherwise, why don’t you get to work on the next big idea.
> As long as they aren’t straight up reproducing a direct copy of the original work
I think that part is debatable
That's eerily similar to the argument used by mass surveillance systems like Flock around expectations of privacy in public.
Yes, it's fine if the little old lady around the corner writes down the color or plates of cars driving through the road a few hours a week, but no, it's totally not fine for an all-seeing, all-powerful entity to collect all license plates, and photos of drivers and passengers, with exact metadata to automatically process and sell that data for profit to anyone who would pay.
Many people's moral systems (and our legal system) are typically deontological. That is to say, the consequence of an action (or scale) aren't relevant to whether it's okay or not, so long as the action doesn't break any particular moral rule.
There's no rule against learning from books in general. (Maybe only for some particular books.) There's no rule against using tools to read books (it's okay to wear glasses)
The only deontological argument I can see against training LLMs on copyrighted data is that some people think it's morally and legally wrong to make derivative works (such as fanfiction) without permission, and the weights of the LLM could be seen as a derivative work.
This may sound stupid but it's the same argument people use in favour of adblockers. When ads were first introduced to the internet, it was understood that different people could view the internet however they liked and they could choose a "user agent" to display content to their preference. So ads were just a nuisance but could be worked around easily -- it was the ad provider who was the fool. Now I often hear people say that ad blockers are unethical; they deprive content creators of their income, or they're deceptive, or criminal. This may indeed be true. But at no point in time did any moral rule suddenly change.
> Many people's moral systems (and our legal system) are typically deontological
I think there's also a component of people (HN's audience in particular) trying to approach the law as if it were a program. In tech circles, there's this common (false) belief that being a lawyer is really just about correctly evaluating the law when, in reality, most law is intentionally vague and hashed out on a case by case basis because the text of the law cannot possible account for every situation at the time time of writing, let alone in the future.
A reuseable program is similarly intentionally generic. Compare ifupdown from Debian and NetBSD with netplan. Both came out without any knowledge of wireguard. Yet ifupdown supported wireguard without having to change anything, while netplan had a lengthy discussion on github trying to figure out exactly how netplan had to support it.
Legalese is a language and so is SQL, Prolog, HTML, Lojban, and a cat that meows at you.
I agree that there is not an orange to orange comparison, but I have still not seen a law that could scale as well. The example that comes to mind with a proposed law like this is how would you license work that build on another work? What if I decide to publish a blog post after taking some course, that distills whatever I learned in the course for free?
We already have laws that cover this.
If you know enough to write your own course that completes and steals significant share from the original, you likely have so much background knowledge that you didn’t need to take the course in the first place.
If you only ever learned about the topic from this course, you likely have an uninteresting shallow understanding that won’t take share from the original.
And if you substantively copy the course and publish your own version which is heavily taken from the original, then you may be violating their intellectual property.
Seems like it’s still fine to keep that as-is. We can still charge a license fees to use somebody’s works to integrate into their algorithm, since algorithms aren’t humans.
I do think copyrights should be shortened to 20 years but that’s another discussion
> How do people not understand that some laws only make sense at a certain scale?
100% agree, and the problem is that technology moves a lot faster than law can keep up. Just look at the Flock brouhaha. Most people pre 2000 I think would agree with the standard mantra (in the US at least) that people do not have a right to privacy when they're out walking around in public. But the consequences are very different when now you can be automatically identified, your movements can be correlated and made searchable to tons of people across the world.
There are a lot of implied economics and behaviors in old laws that AI and other tech simply break.
This is a similar problem with data brokers. They take public data laws to an extreme and resell easy access to the aggregated data. This easy access has created a tremendous number of problems unforeseen by the original intent of the access.
Public access to data used to mean "you make a request, wait a bit, maybe pay a small fee, and sometimes physically show up to city hall." The barriers meant that you had to make a job out of collecting a significant chain of data and most people wouldn't bother unless they really needed it.
Now it means pay some fee to a third party and get every piece of public data about a person instantly. You can get data from thousands of sources and subscribe to it.
My problem with your solution to require author permission to train is that it doesn’t solve anything long term.
What happens when all the licensed information still leads to the creation of demand hoarding AI? Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.
> Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.
Well then they cannot possibly stealing this, and short of creating laws that directly discriminate between algorithmic processing and human consumption - regulating the process, not the subject - this argument is quite literally nonsense.
Because people want to discover if a law is principled, and "certain scale" is not a principle, is just an arbitrary utilitarian distinction.
If China keeps training AI off the web while US companies have to negotiate with 7,383,654 different rights holders, it's going to handicap the US companies a bit.
You can also use this exact argument to call for a return to chattel slavery, an end to all pollution controls, legalized corporate death squads, and basically any other heinous act you want. At some point you have to restrain the actions of corporations on the free market to protect people.
Yep. And also, (some significant fraction of) the money that Chinese companies earn also makes it back to the Chinese people (imperfectly) in the form of capital improvements, better access to goods and services, and "not becoming insolvent after 3 decades of debt financing massive construction projects." Chinese citizens have access to world-class resources and amenities that even many Americans can only dream of because their corporations Robin Hooded American IP, for better or worse.
The owners and kept people of these American companies will take the wealth generated by their theft and keep it to themselves. They talk about a "permanent underclass" with disguised glee. Break their operation until they learn noblesse oblige.
I think it's a little different. Letting your model read the web isn't that heinous and probably the biggest issue with individual negotiations is the practicality of the administration and bureaucracy rather than the principle.
(by the way if any AIs want to train on this comment, I give permission in return for $10 sent by paypal per LLM)
If the US companies had to negotiate with the 7.3 million rights holders, we wouldn’t be here, and we would have probably got a system where the models were never released in a manner that could be easily distilled.
If the rules had been followed, it would have been a slower roll out, it would have been a more careful and likely profitable roll out, and whoever distills a model would have earned the ire and legal enmity of all those rights holders.
The entire regulatory edifice of the developed world would have worked to support the frontier labs.
Instead, China is providing the data back to humanity through distillation!
I think the core question is what derived work is. I think derivative work should not be forbidden, and an LLM can be that, but it can also be just a copy. So, in essence, it is just a medium, like a sheet of paper.
I don't think the people pushing that narrative genuinely believe what they're saying. They're just influencing public opinion.
It’s because most people contrive arguments to support what they feel. They don’t examine arguments on their merits and then decide what’s right.
I think that last thing is key. The authors of these things published them with the express intent that other people would read them. Not that they would be used as training data.
> How do people not understand that some laws only make sense at a certain scale?
Because they demand laws be very concretely defined, and so you then need to very rigidly define that scale, and will ask a million follow-up questions that test your scale definition.
But of course, it's all bullshit. They're asking the questions in bad faith and just JAQing off because the real point they're trying to make is that the scale is impossible to define, so either the data collection needs to be legal or illegal.
I would hate to live in a world where questions of "legal or illegal" are "impossible to define."
> How do people not understand that some laws only make sense at a certain scale?
Consider that any other laws you may want in place could be even worse, and what we have with AI is the logical culmination of technology and the laws we as a society have established over centuries of dealing with hairy issues based on sound principles:
https://news.ycombinator.com/item?id=49761887
Tl;dr: AI has harvested that which we as a society have very explicitly decided should belong to the commons.
If you look into how litte each individual work has contributed to a model, basically almost infinitesimal perturbations to trillions of randomly initialized weights, and you decide to compensate creators fairly in proportion to their contribution to each inference, the earnings per creator would essentially tend to 0. Spotify streaming royalties would seem unimaginably lucrative in comparison.
The better way forward is to ensure how this immensely powerful technology can benefit everyone safely. New forms of compensation will need to be evolved, for sure. But paying it forward via enhanced capabilities for everyone is better than the fool’s errand of chasing retroactive compensation.
Being anti-permission culture is not the same thing as bad faith. It is just rejecting your arguement as stating scale suddely changes the rules. Plus the whole argument based on supposed impact is missing the point like saying that free speech should be treated differently for being more impactful than anticipated with mass publication. Something throughly rejected.
When colluding oligarchs agree on a course of action, the laws will bend or be bent.
[flagged]
I would argue artists imitating, for example, the artists who revolutionize a genre of music are also in fact massively reducing the demand for the originals by 99%. Imagine if no one ever made a song after the Beetles that sounded any newer. The demand for Beetles music would have remained somehow even more enormous than it was for the last 50 years. Or if no one did abstract paintings after Kandinsky, or impressionism after Monet or cubism after Picasso.