It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.

I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.

There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.

The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):

From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...

> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”

Can they keep backup of the digital copy?

It's a good question.

This site, https://copyrightalliance.org/education/copyright-law-explai..., states "It is important to note that this exception for backup copies only applies to computer programs and not to other copyrighted works, such as digital movies, music, or photographs or ebooks." But it seems unbelievable to me that they would have spent millions copying all these books and not have backups.

> I doubt AI companies would use a single physical book if they could avoid it

They just don't want to pay what the copyright holders want to charge

Ai companies don't use ebooks, because they are more expensive than second hand books

They absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.

The entire ironic thing here is that a huge part of those 81 terabytes of ebooks that Meta torrented were directly pirated books from Anna's Archive.

I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks

On Amazon right now, retail prices for e-book copies are higher than for the corresponding paperbacks.

This is exactly why Amazon has also been doing this acquisition and destructive scanning of millions of books for some time now.

i imagine its because the doctrine of first sale does not apply to ebooks.