This settlement has basically nothing to do with LLMs.

At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.

But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

Where Anthropic ran into problems is that they put all their pirated books into a big central library (file on a server), and planned to keep those copies forever. Including copies they never actually fed into the LLM (a point that seriously worked against them).

Alsup ruled this central library of pirated books was copyright infringement. And it's this "pirated central library" that Anthropic are now paying a a 1.5B settlement for, nothing else.

The fact that the pirated books were also used to train LLMs is legally irrelevant. Though... I suspect a non AI company could have negotiated a significantly smaller settlement.

[0] https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...

As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind China, who doesn't give a shit about intellectual property, but that doesn't mean we aren't crossing an ethical boundary, acting like them.

> As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use.

It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.

Yes, ultimately the problem is that the law is vague or inadequate. The courts have their definitions of fair use, which are their best efforts at interpreting the law, and I have mine, which is different.

William Roper: "So, now you give the Devil the benefit of law!"

Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?"

William Roper: "Yes, I’d cut down every law in England to do that!"

Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, from coast to coast, Man’s laws, not God’s! And if you cut them down, and you’re just the man to do it, do you really think you could stand upright in the winds that would blow then? Yes, I’d give the Devil benefit of law, for my own safety’s sake!"

This is why the idea of being "Vogelfrei" or "lawless" was honestly a terrifying concept in the middle ages. They are neither bound by law, nor protected by law.

A lawless man can be struck down with force without persecution by law, because they are lawless.

It's not a settled area of law and there is a SDNY judge that has a completely different application of the fair use analysis in the same exact context and came to a completely different conclusion (that it is not fair use).

I would like to see a citation on that b/c I am unaware of it. The only case I see in SDNY is the NYT v OpenAI case which has not been ruled on yet. https://www.reuters.com/legal/legalindustry/copyright-law-20...

Sorry, I'm thinking of Kadrey, where the court rejected Anthropic's "training" argument and provided an explanation as to how author litigants should demonstrate market harm in order to succeed on a fair use analysis, a factor that Alsup did not effectively weigh.

I suspect the market harm angle is not going to work out either based on the one study I know of on the topic: https://www.nber.org/papers/w34777

> We document a tripling in the number of new books coming to market between late 2022 and late 2025 that mirrors the use of AI that we detect in new books. The effects of this influx on consumer welfare depend on the quality of the additional books. The average quality of new books has fallen with the LLM-induced influx, and books with detected AI are substantially worse than human-authored books, so that much of the new work is of little value to consumers. Still, the LLM influx has delivered some books in the middle range of the usage/quality distribution, and the LLM-era entry process delivered seven percent more consumer surplus from books than the pre-LLM process in 2025.

...

Moreover, the arrival of LLMs does not appear to have displaced activity by incumbent authors. Despite the controversy surrounding LLMs, their effect on book consumers – like other cost-reducing technological changes in the cultural industries – is positive. However, because the new books are mostly of low quality, the effects are modest

So not only are existing authors unharmed (because most of the new competition is slop) there is even a small improvement for consumers.

What do you know, you can get someone to support any message or arguments you want.

Lesson in there about experts and politics.

IMO, using copyrighted works to train models should only be "fair use", if the models are then released as (at least) open weight, so that the public can benefit from it. (Although as noted by a sibling, this would require a law change, not action by the court).

I’m not sure if that’s enough but it would be a great start.

Adobe’s ereaders had a disclaimer that their books cannot be read aloud. There’s clearly precedent that this sort of transformation was disallowed by publishers at the time. Interestingly, at least the audiobook of the latest dungeon crawler Carl has a disclaimer that it can’t be used to train AI

"I haven't loaded an advertisement in 20 years, I have 6TB of movies, 2TB of music, and seemingly endless file trees of mangas, all acquired for free over the years. Now having not said that, I beg you enforce copyright on these AI labs, so I can get a cut of their revenue for my years of writing well researched comments on the internet"

The internet, in true internet fashion, still has the general logic level of a 15 year old.

[deleted]

In that case AI should just be open source/weight. I don't agree with copyright in general but I see where you're coming from.

I think this posture is hugely beneficial to China if they can commoditise the hardware.

So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything.

I’m guessing you have some kind of imagined idea of some small author being compensated handsomely for his book and all future earnings that could have come from it. Reality though is that between the attorneys that will run away with some high triple digit millions and the corporations that own the rights to the subject works, there will be measly “checks” for any actual person that created anything, i.e., an artist or author.

In an odd way, this whole case is really just “capitalism” cannibalizing itself, i.e., publishers greedily and also in a terrified manner trying to steal away as much capital from the technological shift to AI as possible in order to either create a buffer or fund their transformation to adapt to what AI means to the very nature of writing itself, let alone publishing.

I suspect human writing could survive, but I don’t see any room for publishers.

> So will you owe life long compensation for all the knowledge you got from books too?

No because we are people and the laws differ for people, corporations, and machines.

I think people forget that laws are perfectly capable of carving out exceptions, leaving purposeful ambiguity, expressing intent, etc. Yes, humans can have special rules, and very obviously should since laws exist to improve human lives.

I am actually not settled on either side of the matter and I have not forgotten that, but I think what we are really looking at is a rather more complex matter than people want to make it out to be. We are holding several but at the very least contradictory positions and they are incompatible.

Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?

Of course exceptions can be carved out, but they cannot be just, inherently. The problem is that we have allowed our ruling maniacs to create a fiction that organizations are people, which not only have more rights, and less responsibilities, and even less consequences/penalties; but also confers upon the individuals that make up the corporate person rather extreme super powers like being able to commit crimes up to outright murder, and there not only are effectively zero consequences for or to them but in most cases today they immensely profit from it and then shield that money from the victims seeking justice.

The underlying issue, why I am not settled on this matter, is that it is inherently contradictory because the facts and underlying assumptions are all so distorted and perverted that there is no good answer to be had and it's really just a matter of rule of power, feigning rule of law.

> So will you owe life long compensation for all the knowledge you got from books too?

You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copyright law either, because there's no copy.

Yes, for some texts that's possible. But for the vast majority, it is not.

> spitting out verbatim text

The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.

The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...

https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...

It's possible. Would be interesting to see their evidence, and to know whether they can reproduce it for arbitrary articles, not just ones that have been endlessly republished on the net.

> spitting out verbatim text

> had produced near-verbatim replicas

Don't get too hung up on the preciseness of the copy - the courts won't. I doubt that spitting out an existing article with a few adjectives changed would be considered transformative.

> if you can't persuade the machine to spit substantially the same text back out verbatim

That's exactly what they've done in a number of the lawsuits, so I'm not sure why you think that hasn't occurred.

Can you cite any information on this not being possible for the vast majority?

Or is it simply that the correct prompt hasn't been written for all possible cases?

I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to.

It very clearly is still compressing the information into the vector weights, and then recovering that information, thus the information is encoded.

Why is a vector database somehow completely different from maintaining a library of the text itself?

Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.

Is that relevant? I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright.

LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.

> but that derived image would surely be under copyright.

I wouldn't bet on that. https://en.wikipedia.org/wiki/Campbell%27s_Soup_Cans

I don't know that this really challenges anything relating to compression.

Why would a derived image be under copyright?

Because that's legally the case? I don't understand the question. Using a lossy compression algorithm on an image does not remove its copyright protection.

Nah, I can't prove a negative. But Common Crawl is 12 petabytes and is not the largest part of what these models get trained on. DeepSeek v4 Pro is, what, 865GB?

That's one hell of a compression ratio, if it can do what you claim.

You are asking to prove a negative. But even assuming that the model is capable of returning every bit of its training data verbatim (a mathematical impossibility) that would not be enough as mere capability is insufficient here. If capability alone were the standard any library that also has a photocopier / scanner would be in violation.

To prove distribution of copyrighted materials it would have to be practical and actually used in the wild by people to circumvent copyright and generate copies of those works. Again, I can't prove a negative, but that isn't the standard, and nobody has shown a practical exploit here.

There was a paper a while back where (from memory) they managed to coax 75% of the original text of some internet-popular books out of an LLM. Harry Potter, 1984, etc. That's why I said it was possible for some texts.

My assumption is that multiple copies in the training data "wear a deeper groove". I believe those are infringing, and should be dealt with on a case-by-case basis. But the vast majority of text doesn't wear that groove.

(Edit: Think it was this one https://arxiv.org/abs/2601.02671)

This reliably inevitable rationalization comes up in every thread it seems, and its ultimate goal is to humanize AI. This is what the big guys want us peons to believe and it works so well, I have even been lectured by an AI for being rude, the implication was that I was logged in and it would be a shame if anything happened to my account.

Quit trying to make AIs human, people who are trying to make AI human keep forgetting that humanized AI's have only the morals relevant to their mission, there is no profit in humanizing AI's because if we continue on this track of humanizing AI's, we being stupid humans will grant them civil rights expecting these new AI's with rights will somehow respect our rights and thats a fundamental misunderstanding of how AI'S actually work.

copyright is bullshit

Copyright is what stops someone from copy+pasting a book that took years to write, then selling it $1 cheaper than the original author on Amazon or whatever and making a margin 1 million percent higher than the original author.

Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…

How many times were hugely popular books rejected before a publisher decided they were worthy?

Copyright far more protects the wealthy than the good. They don't need to sell your book, they just need to own the book that people are buying right now. Giving your book a chance to sell would dectract from those sales.

If there were no copyright anyone trying to sell the book $1 cheaper would be undercut by someone selling $1 cheaper them them, and so on. The financial incentive to do that goes away. People then choose to distribute based on different incentives, like the fact that they have seen something worthy that others should see. We have almost completely lost that today because the financial incentive doesn't care what it is as long as you buy it. That might lead to a world dominated by an optimisation for whatever it takes to get you engaged, or worse, addicted. That world might really suck.

There needs to be a way to support the creation of art. Copyright lets a few corporations decide the subset of available art is seen enough and available to pay for (in the hope that maybe some of the patment gets to the creator). It is not a system that works in the modern world.

Well, there is nothing to distribute if the author is not incentivized to write... which you seemed to skip past.

Indeed, perhaps I should have said

There needs to be a way to support the creation of art.

Great, that way is called copyright. The author has the right to control who has the rights to distribute their work, and can require compensation in exchange for that right; what economists refer to as "selling".

People have been paid artists before copyright even existed, what some might call patronage.

And?

Therefore copyright is not necessary.

That's a big leap

>There needs to be a way to support the creation of art.

Yeah, it's called "copyright."

If money is the only incentive, then it's not a product of artistic work.

Also current copyright laws only exists to fulfill the constitutional mandate to promote the progress of science and useful arts. There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.

>If money is the only incentive, then it's not a product of artistic work.

This is just bullshit and no one said it's the only incentive.

> There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.

such as??

> Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…

Yes, all the rich class propaganda being pushed by open source developers working on software in their free time.

That's oversimplifying things.

Copyright doesn't actually stop me from pirating a book or an mp3 right now. Heck, I'll just download a book right now. Bam. Done. Some things are so difficult to keep from being pirated, such a photographs, that saying the copyright system protects photographers strikes me as a bit silly. It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.

Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption. In fact the "intellegence is a ultility" ramblings of Sam Altmen sort-of point in this direction. But that's just one idea thought up early in the morning when its too hot to sleep properly. I'm sure there are many others.

> It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.

That is a very small slice thanks to copyrights. Without copyrights then corporations stealing from the small guy like this would be the majority of it.

It's funny because copyright only benefits the rich now. Record labels hold all the copyright to songs, same with publishers for books, Disney made sure it lasts over a hundred years. The days of copyright being held by individuals in any real sense is long gone.

Right, it's about incentivising intellectual work. While I have big issues with the copyright system, like all the extensions lobbied for by Disney and friends, it did enable a lot of good work to happen.

> it did enable a lot of good work to happen.

How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system?

A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?

Yeah, I can't AB test against a world without copyright at all, but I think there's sufficient evidence to believe that a lot of stuff would never have gotten done without copyright to ensure it could be done gainfully.

The importance of striking a balance between incentivising creation and enriching culture was why the original copyright term was dramatically shorter. The modern term of owners life + 80 years or whatever it is, is clearly ridiculous. 20 years before entering public domain seems pretty reasonable.

There's unfortunately also some pressure against people using legitimate public domain works. E.g. youtubers getting copyright strikes for playing public domain music because it's too similar to a specific copyrighted recording.

You’re arguing that freely remixing original work will give rise to greatness that’s even better than original work?

Isn't most creative work synthesis rather than unique whole-cloth creation? Look at what happens with software when it is open sourced and allowed to be remixed freely. Are we better or worse off because of it?

There's nothing that prevents people from remixing things that are not copyrighted and create something amazing that others are interested in or of cultural value.

With open source, I should note, its remixing is in fact governed by copyright.

Possibly. Sort of like how Disney remixed basically everything from existing fairy tales for decades then made it so nobody else could.

Yes? I don't understand how this is even a question, this is exactly how it worked throughout human history.

In what way is copyright preventing it today?

I can't make and sell a remix of an existing IP and even in cases of free distribution it can get dicey legally.

I think the idea is if "AI" can solve math proofs that humans haven't for a century then if "AI" freestyles stolen art and literature then it might create something as good if not better because of resources and processing power

There might be something to that logic but art and literature doesn't obey rules like math and copyright exists to protect creators

Go read a few fanfics and tell me you still think that theres added value.

Quite rich coming from a creator apparently, only certain creations are deemed worthy by you, seems like it invalidates your entire stance on the sanctity of art that all these pro-copyright people seem to hold.

I can't tell if this comment is satire or not.

You're speaking to the generation of pirates. What? Suddenly everyone is hanging up their high seas hat to capture the virtue signals of current sentiment?

I would be fine with abandoning copyright ... If it is done for everyone equally, and not just tech giants and VC money businesses get a free pass, while everyone else still has to follow the copyright laws. Lets go ahead and usher in an age of free information and experiencing all forms of human expression for everyone. But lets also come up with a way, to compensate our creative minds and our educators and artists. How about that UBI? We stand much to gain as humanity.

This. Copyright is a flawed system. There can be alternatives that allow more than 1 player to play and not create monopolies.

For example. I invent a new method of power washing. I start a power washing business using new tech. I file the tech for patent and copyright-equivalent use. This is then made available to other power wash companies that wish to use the tech and be certified in it so long as a small portion of their revenue goes back to the inventor for a set amount per volume, or something similar of a metric that has a cutoff after a point.

This will breed new industries, create new jobs, introduce new innovations, and allow the markets to move on from being strangled by one giant corporation.

Isn't that just...patent licensing? But I agree that it should be a forced outcome so everyone can use it rather than waiting a ridiculous 20 years.

This reminds me of this argument with libertarians/ancaps:

A: rich people pay less % in taxes than wage workers, we should close the loopholes

B: but taxes are immoral to begin with

A: ok, but can we do something now about the unequal enforcement? Unrealized gains, tax havens, trusts, fake charities, etc?

B: well a society based on property rights… ackhully you should read this book by Mises/Rothbard/Rand

Never really understood how libertarians expect to have someone making guns for their fiefdoms when there is no one to enforce property rights for said gun elements and manufactories.

Libertarianism is not a philosophy. It's selfishness taken to extremes and trying to find ways to justify it at a societal level. The only reason we're the top species is because we're ultra social and have culture, which is inherently a social trait (don't eat those red berries, they're poisonous). Libertarianism want all the benefits of working together with no actual thought into how that working together happens in real life, including punishment for bad behavior.

Yeah its wrong in very very similar ways to communism.

What very very similar ways would that be?

I guess they presume it requires on the good will of everyone to live peacefully without violating (non-existent) property laws over, say, robbing you in your sleep.

[deleted]

Surprise: the comment above was downvoted in the bastion of libertarianism :-)

https://youtu.be/lh2__MN-FTU?si=LXIaljh__s8fD75l&t=1568

About 3 minutes of video worth watching.

> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

The court says otherwise.

> Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.

Then it says it doesn't need to decide on that basis because they kept it not just for training LLMs, but also for building a central library. Which seems a bit ridiculous, because the sole purpose of the central library is to train LLMs.

> At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.

> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they eventually deleted them afterwards)

The way I understood it, was that essentially the entire case rested on if Anthropics use was "transformative" or not. And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.

Regardless if they deleted files or not, if nothing existing was transformed, it would have been illegal. But because of the destruction of k̶n̶o̶w̶l̶e̶d̶g̶e̶ physical property, this ended up being legal.

> And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.

You have to be careful, just because the judge points a factor out as notable, doesn't mean that factor was required.

The destruction of source books makes Anthropic's fair use argument [2] especially air tight, but it would be a mistake to assume that act was required, or is what made it transformative.

In the previous google books case [1] (which this case cites), google borrowed books from libraries, scanned them, then returned them. They were not destroyed, google didn't even keep the physical copy.

Yet Google Books was ruled fair use, because it was transformative.

[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

[2] Note... This part of the ruling is still not about LLMs. This was about Anthropic's right to scan books and then keep a digital library of them.

What does that mean to be "transformative", as a defense?

I thought that was explicitly disallowed use... like turning someone else's book into an audiobook and selling streaming access to it.

Or writing a film adaptation and selling the film.

Clearly I was thinking about it all wrong. Those wouldn't be allowed, even if you legally aquire the book from a store or library.

Transformative alone isn't enough for a fair use defence. Nor is it required. It's simply one of the many factors a judge will take into account.

But it was an important factor in the google books case.

One of the other key factors is how it impacts potential sales of the original work. Turning it into an audiobook might be transformative, but when you sell access to it people will buy your audiobook instead of the original book. So it's almost certainly not fair use.

In the google books case, google scanned the books but didn't distribute the content of the books to the user. They only distributed the transformed ability to search books to users. The sales of the books weren't impacted negatively, because the user still had to acquire a copy of the book from somewhere else if they wanted to read the whole work. In fact, google books arguable increases sales of the original work in some circumstances.

My understanding comes from here, seems pretty clear to me but won't claim to be a lawyer of course:

> Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position.

https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

Based on that I get the impression it's quite literally the destruction part that makes it transformative, without it, it wouldn't have been tranformative at all.

I've read through the order again. I can't find anywhere where Alsup says the destruction was required.

He cites three cases where a conversion from one format to another (without destruction of the previous version) was ruled to be fair use. Including scanning books with the google books case. (And referenced the Napster case, where a similar argument was rejected)

Then made the following comparison.

"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."

So it wasn't transformative because of the destruction. The destruction only made it "even more clearly transformative" than those other cases.

Like, how can destruction be required if there were previous cases where it wasn't?

The key legal point is not that Anthropic destroyed the books, but the key fact was that Anthropic didn't distribute the scanned copies. Alsup keeps returning to this point:

"But what matters most is whether the format change exploits anything the Copyright Act reserves to the copyright owner. Anthropic already had purchased permanent library copies (print ones). It did not create new copies to share or sell outside"

"But again, the replacement copy here was kept in the central library, not distributed"

The conclusion of that section doesn't even mention the destruction at all.

arstechnica isn't exactly wrong, the quote also mentioned "and kept the digital files internally rather than distributing them". It just put way too much emphasis on the destruction, and not enough on the lack of distribution.

The other thing that arstechnica are missing:

Antropic didn't destroy the books because they thought it would strengthen their legal argument. They destroyed the because it's a lot cheaper and faster to scan books by ripping off their bindings and feeding the stacks of loose pages into a document scanner.

> So it wasn't transformative because of the destruction

I mean, the parts of "in order to save storage space" and "The print original was destroyed. One replaced the other." again makes it clear (to me at least) that the destruction is pretty much what sticks out here that makes it "more transformative" (whatever that means) than the previous cited cases.

But yeah, agree that also "didn't distribute the scanned copies" seems to have mattered a great deal, as well as the destruction part.

Nope, that's a little bit of sloppy writing on the part of Ars. I am not a lawyer, but I'll be happy to discuss the technicalities with anybody here. I'm fairly passionate about the technicalities of copyright.

Yet countless families, including old folks were ruined during untold numbers of RIAA suits because "converting to save space" is not a permissable use.

They used to go around destroying lives by the thousands after Napster was creating because of the invalidity of that argument.

It is a crime to make a CD of your MP3s and vice versa, and you cannot convert your VHS to DVD.

A billionaire does it at scale, well then saving space via format conversion is a grand, while the peons still can see their lives destroyed but with it hidden via the CCB secret panel. Two tier American Justice on full display. Bankrupty and seizure or worse for thee and billions for he. Format conversion legalized only for oligarchs, and of course, no appeal so it will only be a binding precedent on that one rich guy and nobody else. Tribe on both sides, keeping special rights for themselves that are illegal for everybody else.

This is not backed up by any evidence. Ripping CDs was never illegal. The DMCA made the circumvention of an effective copyright protection mechanism illegal, which made ripping DVDs and Blu-rays a crime. But that's separate from copyright itself. The RIAA sued Napster users not because they were converting files, but because they were obtaining them from others without a license.

There is a lot of evidence. I lived through it. Every family with children and an internet connection or MP3 player was terrified of getting ruined suddenly via a letter. It was in the news every day about some other grandpa or single mother losing their house.

Ripping CDs was long illegal. Perhaps the Librarian of Congress made an exception. Now they hid everything behind a CCB that is like Arbitration so we will never know because they have hid almost aspects of societal justice about copyright and business labor behind arbitration style secrecy. The most useful courts are secret and now people believe there are no proceedings and they do not understand how much of our society was litigated and debated before.

Here is an article from 2008 Specifically explaining that ripping a CD to format convert for personal use is illegal and the RIAA and Sony BMG saying it merited suit but they had bigger fish to fry.

1. https://www.npr.org/transcripts/17814972

Mr. FISHER: That's right. So then, you have to ask yourself, why is the industry continuing to cling to that notion that there is no such legal right? (1)

"Bigger fish to fry" is not the same as "legal".

People selling software to easily convert VHS to hard drive were also punished. For decades they were very clear that format shifting was outlawed. But now that it supports centralizing power and creating a permanent class of info-priests to rule the society, they allow it for them.

Frankly, making all the justice system secret is why the media had to turn to personality cult nonsense for most reporting. All the great stories of the past were informed via the justice system activities. Since all the court stuff is secret now, all they had to talk about was Donald Trump.

"Discovery" provided the bulk of news facts before they secreted away all the justice system proceedings for liability, labor, negligence, medical care, copyright.

It used to be possible to know stuff about America and there was "evidence" all over the place. Now there is never any evidence for anything anywhere. That's Scalia's legacy thanks to Concepcion, absolutely gutting the ability of the society to use Hawthorne effects to discern legality and behavior.

This country used to have evidence for everything, and now a lack of evidence is so common that it is a trope level popular refrain.

You're confused about what was actually illegal and what the industry wanted you to believe was illegal. They didn't want to take anybody to court for actually ripping a CD because they didn't want to lose and have the precedent set like it was in Sony versus Betamax. This isn't "bigger fish to fry". This is "terrified of the precedent".

> The dispute arises from a suit the RIAA filed against a man in Arizona who bought CDs, copied them into his computer as MP3 files, and then put them into a shared folder that other people could access through Kazaa, a computer program for sharing music. He's being sued for that last part.

Then they state what they wish were true:

> But according to Marc Fisher, legal documents and some statements by industry officials make it clear that the industry regards the simple act of copying a CD onto your computer or your iPod as illegal.

But just because they wished it to be did not make it so. Trillion-dollar companies have provided end-users with software to rip CDs (including iTunes), and there's never been a court case over it.

I'm not aware of any cases where the RIAA sued people who ripped their own CD/DVDs/VHS for personal use.

Their MO was suing owners of internet connections which were seen sharing content on file sharing networks.

Why is scanning a book transformative(a la Google) but reading data off a CD and putting into a digital format not?

How is streaming bits of the music from your computer not transformative?

The RIAA (and the wider copyright industry) were careful to never bring a case that might rule on the issue of "ripping data from CDs and converting to digital".

What they did was bring a case against Naspter, which ruled that ripping data off CDs AND THEN sharing it to millions of people over the internet was infringement. Not because of the ripping, but because of the sharing. The RIAA then somehow managed to twist public discourse to interpet the ruling as "ripping CDs is illegal".

They were careful, because the Sony Betamax case had already ruled that recording TV of the airwaves was legal, which is already a weaker case than ripping CDs you own. They knew such a case would likely rule against them, and they found the ambiguity to be much more useful.

And later cases like the google books case, and this Anthropic one provide even more evidence that the courts would likely rule that ripping CDs was legal if such a case was ever bought. (Though, it really depends on what you do with the digital copy)

it was a crime to run your own unregistered taxi service in many places until uber came and the laws changed to adapt

They did not change the laws. The rich tribe guy ignored them and they let him off just like the Anthropic guy. They did not change the laws. The old businesses just folded and the new ones via unlicensed independent contractors made cottage industries out of small scale fraud and tax evasion.

How many Uber drivers can show their local business license for every town they pick people up in? How many have sales tax accounts for their state? Every uber driver without them should have been charged with the same crime as Al Capone.

Now they are trying to control the knowledge, and the vehicle driving, etc. via AI.

It is frankly a tribe takeover via mass criminal activity.

He should have been charged with tax evasion for every pickup in a place where he lacked a business license, but the tribe would never allow it. Compliance is only for the other guys.

Only the dumb local guy graduating high school trying to earn a living has to worry about legal compliance since the rich guys are too hard to prosecute.

They did not change laws. They refuse to enforce them and we are being taken over by the reincarnated legion of Al Capone as a result.

> They did not change the laws.

yes they did?

in Québec: https://www.ctvnews.ca/montreal/article/uber-is-officially-a...

in France: loi Thévenoud and Grandguillaume (which were the follow up to negociations between the french gov' and uber), etc.

other countries are the same around the same period, e.g. https://legislation.nsw.gov.au/view/html/inforce/current/act... etc etc

You're mistaken.

The transformativeness of the use is independent of the destruction of the books. The destruction of the books allowed them to argue that they had not duplicated them, and was instrumental in the argument supporting the legality of scanning them. But that's entirely upstream of the way the data was leveraged, which is what is critical in the argument about the use being transformative.

That's a good ruling, because otherwise only the big companies can afford to pay for enough content to make an LLM (say goodbye to open weight or research LLMs). Having a fee like this is actually a form of regulatory capture.

>But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

You're not. Even if training is fair use, it doesn't mean you can steal copies to train the model. It just means the training itself isn't an infringement (in Alsup's opinion). Stealing the copies of the books was an infringement and that's exactly the liability that Anthropic settled.

[flagged]

Who would have thought, that this is the way, which we take to arrive at the burning books stage again? They neatly line up with historical perpetrators in that regard.

It's easier to ask for forgiveness than permission, right?

It seems to be the modus operandi of corporations in general: they commit any kind of infringement they want and then later they go for a settlement with a value that's, of course, not too big for a company too big to fail.

In the meantime, the average person or company gets shafted.

In my opinion, we are one step away from AI companies capturing the entirety of copyright legislation.

You do indeed appear to have a valid point. Many "chosen" companies, like Uber for example, appear to have broken numerous laws. Legal action against many such companies comes suspiciously slowly, where they have already obtained massive profits and value, before the possibility of being shut down comes. Then, when they are finally pulled into court, they have all kinds of money for the best lawyers and have already paid the right politicians (and others).

When the legal judgements for wrongdoing are finally handed out, they often come across as just an inconvenience or kind of tax, which is easily handled in comparison to the profits they've already made. Yet, if average Joe or persons not considered as being of "the right type" were to do such actions, they quickly get the full book thrown at them. Often, the full measure of legal punishment, where their company and life is or about nearly over.

Paying this sort of fee in the first place is itself regulatory capture because only the big companies will be able to pay it. If they can pirate to make an LLM then so should us commoners be able to too.

Hey I just came up with this idea, I'm going to feed copyrighted books into my LLM that remembers them verbatim, and then people pay me to ask the LLM for complete copies of a book.

Wait, no, not verbatim. It transforms upper case into lower case and vice versa.