I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.

Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.

Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.

> Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.

It seems unreasonable to stop there though; the agentic bots are designed and marketed as able to compete with the initially-scraped sources.

I'm not convinced that a competitive use at one remove should be treated as not competitive.

I think that's more true in image generation than in text? At least, all the money is in LLMs that write code, not LLMs that write O'Reilley-style coding books.

If you have a websites that offers guides, how-tos or tutorials, LLMs directly compete with you. StackOverflow would also have a really good case

After all LLMs don't just code, they also answer questions and give step-by-step instructions. In terms of total userbase those features are used a lot more than writing code

I thought LLM-authored books were rampant in Amazon, though?

I think LLMs providers pretty directly compete with content they scrape like Wikipedia and SO...

Remember the golden rule of the golden rules:

Who has the gold makes the rules.

If a society operates under a rule like this, it is no better than Russia or any other tyranny where law is for me but not for thee. This is not how it should work in a supposedly free and lawful country.

That’s how works everywhere

You're downvoted, but it's true. Some places just mask it with fancy processes and manners a bit better.

It’s been working like that in the US for a while.

Do you think the bottom 99% of the country would ever win a legal fight with one of the tech billionaires?

Even if they were 100% in the right, they could just drag out the legal process with endless motions and appeals until the average Joe lacked the funds to continue the fight.

Well I have heard that some people managed to sue some corporations and actually win some cases before.

PFAS case is a rare occurrence, and this is after DuPont knowing the harms and downplaying or ignoring it for decades.

A similar thing has been surfaced in Dalton, Georgia, again about PFAS [0]. Let's see how will it play out...

[0]: https://www.pbs.org/wgbh/frontline/article/pfas-forever-chem...

When I said the tech billionaires I meant the actual oligarchs.

Companies will eventually take a cost benefit evaluation and stop fighting if the risk is too high or the fines too large.

A personal billionaire who decides to fuck you over is capable of behaving irrationally and just beating you via resources.

Look at trumps infinite appeal strategy which apparently works.

He only just was forced to pay Jean Carrol a few months ago, four years after he lost his civil case.

If you’re a regular Joe and some billionaire harmed you in a way that hurt your income. You’re not gonna be able to afford 4 years of legal battles.

As long as money buys power in our legal system, you can’t compete with someone with effectively infinite more wealth than you.

No kidding. I dunno what kind of traffic loss Wikipedia has had but SO is really dead these days

If they weren't competing with AI then why is AI killing it?

> weren't competing with AI then why is AI killing it?

But it is not the AI who is killing it, users are doing it.

There's a difference between creating a market for something better, so that nobody wants the old thing, and competing _in_ the market for the old thing by copying it directly.

If LLMs only made SO redundant by writing code and solving my technical problems autonomously so I never have to think about it, I would agree. But often I do ask LLMs technical questions, and they answer in great detail. And that part is a very direct SO competitor

And what would be a read-only version of X like XCancel compete against, exactly? Ads impressions? That would be the only possible thing yet they don't add any ads.

It's depriving X of impressions that they could monetise, no? Xcancel doesn't have to make money itself, it just has to impair the rights of the copyright holder. Otherwise piracy would also be legal as long as it were non-profit...

> Otherwise piracy would also be legal as long as it were non-profit...

Which is in a few jurisdictions, or at least is not prosecuted if it's for personal use. Also, according to your definition, the creator of uBlock Origin or any other adblock system should be sued in the same way, because they are depriving $ADS_CORP of their precious impressions.

Well, adblockers don't copy the copyrighted content. They just control how it's rendered on the user's machine. Copyright cares about making copies and especially distributing them.

Xcancel isn't copying the copyrighted content either. It takes the raw JSON/HTML data from twitter/x via reverse proxy and redisplays it on xcancel, as opposed to twitter/x. How is that copying?

[deleted]

You have a point on this, I recognize, but it still seems a very thin line to walk (for X) - at least morally, because I don't think they are actually loosing real big money to anyone.

X doesnt own msot of those copyrights, except where its elon musk's own posts.

i dont think the actual copyright owners care, given they put their content onto a vaguely public view where they aim to get the most traffic to something else they are doing

If two things are competing, they are in the same market by definition.

I don't really buy this.

It's like saying "toaster oven/air fryer combos" don't actually compete with toaster ovens or air fryers because they are creating a market for something better

Of course they complete.

Toaster ovens compete with toasters. Microwaves compete with toaster ovens.

Just because it's not the exact same product doesn't mean it's not competing

Would I be allowed to steal LG's designs for a microwave and make a "superwave" that does laundry and heats food? Would you claim those products don't compete because the superwave is "something better"?

It's not illegal to write a similar book or song to an existing one. Copyright only protects existing works from literal copying (possibly in part).

And derivative works, like AI makes

> Precedent is pretty clear:

What cases are you citing when you say this?

Bartz v Anthropic. Though the plaintiffs did get something, it was because of the piracy to the original works (competing against the legal market for the books), not the use of them to train the LLM.

Bartz is an author though.

Is X claiming ownership of the posts people make because pretty much every single social media site doesn't so they have section 230 protection.

They can just round up some friendly users and sue under their names. Starting with their own corporate accounts?

That would definitely limit damages to strictly those accounts.

I'm not even sure he can use his own account as one of them. The SEC might be pretty friendly to him but I'm not sure that limiting access to a location where material information about Tesla/SpaceX is provided won't become a problem.

But I'm not even sure what damages the accounts are suffering as revenue sharing is going away [1]. With Bartz the damage is a loss of sale. With X the damage is $0 per post to the poster.

There is a newer Original Content Rewards program [2] but it seems to split revenue from X Premium and presumably people that have X Premium are not using XCancel so the damages would be 0.

[1]: https://help.x.com/en/using-x/creator-revenue-sharing

[2]: https://help.x.com/en/using-x/original-content-rewards

These statements would suggest that the precedent is not, in fact, clear

[deleted]

but xcancel is not putting ads or making a commercial product

How clear is it when I google a recipe and get an AI-generated recipe that's clearly derived from the top three results and then placed above those results? That sounds like it's both transformative (in that the recipe created by the AI may not match any one of the scraped recipes perfectly) and also competitive (in that the AI takes page views away from the pages it got the recipes from)

> How clear is it when I google a recipe and get an AI-generated recipe

Recipes can’t be copyrighted

Here’s one discussion about this https://www.nycbar.org/reports/secret-ingredients-how-to-pro...

fair, but recipes aren't the only thing where AI summaries at the top of search results are simultaneously transformative and competitive. Any information that is scraped from a website and then summarized by AI in a way that prevents that website from getting views is both transformative and competitive.

Transformation.

Taking something someone else made and showing it as-is, bypassing their own restrictions: No no.

Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company

twitter didnt make it though. they have a license to it

So in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed?

Because that's stupid. These laws are stupid.

That's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works.

That's stupid, these courts are stupid

It should have nothing to do with storing copies it should have to do with what the models can produce. And it's clear they can produce copyrighted works, they've just been tuned so they don't.

That shouldn't satisfy anyone.

> they can produce copyrighted works, they've just been tuned so they don't.

In other words... they can't. A different one can but this one can't. The court is not stupid, and will consider this fact.

They can't unless the tuner produces a copyrighted work for whatever their purpose would be, because they can step in and "detune" the thing at 3am for a competitive advantage.

This assumes the "tuning" (I assume you mean fine-tuning?) is lossless, which it isn't.

You can produce copyrighted works and have just been tuned not to lol. I fail to see how limiting an ability to comply with the law is any different from just complying with the law?

Proprietary Computer systems shouldn't get the same leeway that humans do. Simple as that.

[deleted]

The point is that you can't steal someone else's content 1:1. But you can use it for a different use (say, display the tweet in an article, then comment on it).

except a bunch of paywalled stuff did end up in training corpora