Thanks Elon, now we know that scraping is illegal! Very good to clarify that for future proceedings against the AI thieves.

Everything is both legal and illegal until a lawsuit happens. Then it collapses depending a little bit on the facts and mostly on who has the better lawyers.

I suspect Nitter's first round with lawyers pointed out that scraping is legal, but now they have been threatened with something else than scraping - Elon claims something else the way Nitter runs is illegal, such as the use of fake accounts to circumvent an access control device (DMCA 1201).

At the end of the day, as an individual or a team or a company, regardless of the statue and case law, you have to perform the calculus on your monetary and legal resources versus your counterparty.

Obviously ungrounded and frivolous cases tend to be easier to defend asymmetrically, but if I was X's legal team, there's no shortage of semi-plauisble claims I could throw at the wall and see what sticks.

As an example of this imbalance in action, BrightData is a 'gray area' company that basically does this exact kind of scraping. They have somehow won against Meta Platforms suing them, and even got X's lawsuit against them for scraping -- identical (?) activity to XCancel -- dismissed.

According to Wikipedia:

> In May 2024, a federal judge dismissed the suit, ruling that Bright Data did not violate X's terms of service or copyright by scraping publicly accessible data.[24] The judge emphasized that such scraping practices are generally legal and that restricting them could lead to information monopolies

But does XCancel have the resources of a company like Bright Data, that's funded and used by companies like Deloitte and Moodys?

If the name Bright Data is ringing a bell to anyone, it’s probably because they are a (the?) primary offender running the LG TV “residential proxy” (botnet)

Which, btw, was the least bad thing about those TVs. I support residential proxying as a means to liberate public data that's being held captive by people like Elon.

And the company's previous name was, ominously... Luminati.

Notice it says "terms of service or copyright". If X's lawyers have any intelligence, they'll have a reason why XCancel is not identical to Bright Data. Perhaps this time, instead of claiming it's a copyright violation, they'll claim it's wire fraud because multiple accounts are used.

I wonder if anti-SLAPP laws could be used to shield nitter: https://en.wikipedia.org/wiki/Strategic_lawsuit_against_publ...

Scraping is legal, but IIRC scraping authenticated contents is not.

Isn't Nitter abusing account sign-in for this?

Anyone can make an account, though, and instantly access that content, so Elon doesn't really have any leg to stand on by claiming they're private. It's not the same as Cambridge Analytica scraping stuff you had to have certain privileges to see by tricking the system into granting those privileges.

Elon unfortunately may have a leg to stand on, because it doesn't functionally matter whether "anyone can make an account and see it". I don't believe any legal ruling to date has been fine with that distinction.

Do you mean this will be the first legal ruling about this particular variation of the issue?

Not true. Creating an account requires agreeing to the ToS. And if the ToS prohibit this kind of usage, then legally, it's not allowed.

See hiQ Labs vs. LinkedIn. hiQ was scraping LinkedIn using accounts, and they were found to be in breach of LinkedIn's terms.

It depends what you mean by "not allowed". Breaking ToS isn't a crime, but you might be sued for damages, but how can they show there were any damages?

It seems pretty easy? Someone viewing it through a mirror isn't going to see ads.

But if they ban the nitter operator surely you can't just bypass that ban legally by making a new X account. That there are bans and the operator has to bypass bans by pretending to be a new person I think that this makes it different.

It's like a club that checks ids to ban people. I wouldn't call that club open to the general public.

Nor would it be illegal to report outside of the club about things people said in the club.

Yes, that's what they went on to point out.

To fix this:

1) Nitter offers a deal: download our browser extension, sign up for Twitter, and we'll give you some kind of perk (Amazon credits, whatever).

2) Browser extension surfs Twitter on the user's behalf, scraping and sending copies to Nitter.

3) We find out how serious the legal system is about prosecuting scraping.

Tall order. The main Nitter instances were using tens of thousands of accounts to evade detection, IIRC. Adding money to the mix creates more legal liability and a paper trail towards people who can be sued.

It's much easier if thousands of people set up their own Nitter instances.

The whole point is not signing up for twitter. If signing up for twitter wasn't a problem, Nitter wouldn't exist in the first place.

The account on its own is worthless - the problem with Twitter is not that people have accounts.

But the fact remains, people continue to sign up for Twitter. As long as that's the case, Nitter should probably try to take advantage of it.

Aurora Store does the same thing to facilitate downloading Play Store apps.

I like this framing, like Schrödinger's cat.

Thin skins over at there at X

I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.

Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.

Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.

> Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.

It seems unreasonable to stop there though; the agentic bots are designed and marketed as able to compete with the initially-scraped sources.

I'm not convinced that a competitive use at one remove should be treated as not competitive.

I think that's more true in image generation than in text? At least, all the money is in LLMs that write code, not LLMs that write O'Reilley-style coding books.

If you have a websites that offers guides, how-tos or tutorials, LLMs directly compete with you. StackOverflow would also have a really good case

After all LLMs don't just code, they also answer questions and give step-by-step instructions. In terms of total userbase those features are used a lot more than writing code

I thought LLM-authored books were rampant in Amazon, though?

I think LLMs providers pretty directly compete with content they scrape like Wikipedia and SO...

Remember the golden rule of the golden rules:

Who has the gold makes the rules.

If a society operates under a rule like this, it is no better than Russia or any other tyranny where law is for me but not for thee. This is not how it should work in a supposedly free and lawful country.

That’s how works everywhere

You're downvoted, but it's true. Some places just mask it with fancy processes and manners a bit better.

It’s been working like that in the US for a while.

Do you think the bottom 99% of the country would ever win a legal fight with one of the tech billionaires?

Even if they were 100% in the right, they could just drag out the legal process with endless motions and appeals until the average Joe lacked the funds to continue the fight.

Well I have heard that some people managed to sue some corporations and actually win some cases before.

PFAS case is a rare occurrence, and this is after DuPont knowing the harms and downplaying or ignoring it for decades.

A similar thing has been surfaced in Dalton, Georgia, again about PFAS [0]. Let's see how will it play out...

[0]: https://www.pbs.org/wgbh/frontline/article/pfas-forever-chem...

When I said the tech billionaires I meant the actual oligarchs.

Companies will eventually take a cost benefit evaluation and stop fighting if the risk is too high or the fines too large.

A personal billionaire who decides to fuck you over is capable of behaving irrationally and just beating you via resources.

Look at trumps infinite appeal strategy which apparently works.

He only just was forced to pay Jean Carrol a few months ago, four years after he lost his civil case.

If you’re a regular Joe and some billionaire harmed you in a way that hurt your income. You’re not gonna be able to afford 4 years of legal battles.

As long as money buys power in our legal system, you can’t compete with someone with effectively infinite more wealth than you.

No kidding. I dunno what kind of traffic loss Wikipedia has had but SO is really dead these days

If they weren't competing with AI then why is AI killing it?

> weren't competing with AI then why is AI killing it?

But it is not the AI who is killing it, users are doing it.

There's a difference between creating a market for something better, so that nobody wants the old thing, and competing _in_ the market for the old thing by copying it directly.

If LLMs only made SO redundant by writing code and solving my technical problems autonomously so I never have to think about it, I would agree. But often I do ask LLMs technical questions, and they answer in great detail. And that part is a very direct SO competitor

And what would be a read-only version of X like XCancel compete against, exactly? Ads impressions? That would be the only possible thing yet they don't add any ads.

It's depriving X of impressions that they could monetise, no? Xcancel doesn't have to make money itself, it just has to impair the rights of the copyright holder. Otherwise piracy would also be legal as long as it were non-profit...

> Otherwise piracy would also be legal as long as it were non-profit...

Which is in a few jurisdictions, or at least is not prosecuted if it's for personal use. Also, according to your definition, the creator of uBlock Origin or any other adblock system should be sued in the same way, because they are depriving $ADS_CORP of their precious impressions.

Well, adblockers don't copy the copyrighted content. They just control how it's rendered on the user's machine. Copyright cares about making copies and especially distributing them.

Xcancel isn't copying the copyrighted content either. It takes the raw JSON/HTML data from twitter/x via reverse proxy and redisplays it on xcancel, as opposed to twitter/x. How is that copying?

[deleted]

You have a point on this, I recognize, but it still seems a very thin line to walk (for X) - at least morally, because I don't think they are actually loosing real big money to anyone.

X doesnt own msot of those copyrights, except where its elon musk's own posts.

i dont think the actual copyright owners care, given they put their content onto a vaguely public view where they aim to get the most traffic to something else they are doing

If two things are competing, they are in the same market by definition.

I don't really buy this.

It's like saying "toaster oven/air fryer combos" don't actually compete with toaster ovens or air fryers because they are creating a market for something better

Of course they complete.

Toaster ovens compete with toasters. Microwaves compete with toaster ovens.

Just because it's not the exact same product doesn't mean it's not competing

Would I be allowed to steal LG's designs for a microwave and make a "superwave" that does laundry and heats food? Would you claim those products don't compete because the superwave is "something better"?

It's not illegal to write a similar book or song to an existing one. Copyright only protects existing works from literal copying (possibly in part).

And derivative works, like AI makes

> Precedent is pretty clear:

What cases are you citing when you say this?

Bartz v Anthropic. Though the plaintiffs did get something, it was because of the piracy to the original works (competing against the legal market for the books), not the use of them to train the LLM.

Bartz is an author though.

Is X claiming ownership of the posts people make because pretty much every single social media site doesn't so they have section 230 protection.

They can just round up some friendly users and sue under their names. Starting with their own corporate accounts?

That would definitely limit damages to strictly those accounts.

I'm not even sure he can use his own account as one of them. The SEC might be pretty friendly to him but I'm not sure that limiting access to a location where material information about Tesla/SpaceX is provided won't become a problem.

But I'm not even sure what damages the accounts are suffering as revenue sharing is going away [1]. With Bartz the damage is a loss of sale. With X the damage is $0 per post to the poster.

There is a newer Original Content Rewards program [2] but it seems to split revenue from X Premium and presumably people that have X Premium are not using XCancel so the damages would be 0.

[1]: https://help.x.com/en/using-x/creator-revenue-sharing

[2]: https://help.x.com/en/using-x/original-content-rewards

These statements would suggest that the precedent is not, in fact, clear

[deleted]

but xcancel is not putting ads or making a commercial product

How clear is it when I google a recipe and get an AI-generated recipe that's clearly derived from the top three results and then placed above those results? That sounds like it's both transformative (in that the recipe created by the AI may not match any one of the scraped recipes perfectly) and also competitive (in that the AI takes page views away from the pages it got the recipes from)

> How clear is it when I google a recipe and get an AI-generated recipe

Recipes can’t be copyrighted

Here’s one discussion about this https://www.nycbar.org/reports/secret-ingredients-how-to-pro...

fair, but recipes aren't the only thing where AI summaries at the top of search results are simultaneously transformative and competitive. Any information that is scraped from a website and then summarized by AI in a way that prevents that website from getting views is both transformative and competitive.

Transformation.

Taking something someone else made and showing it as-is, bypassing their own restrictions: No no.

Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company

twitter didnt make it though. they have a license to it

So in theory if you took twitter content and then transformed it so it "summarizes" all tweets with an AI rather than posting the exact text, would that be allowed?

Because that's stupid. These laws are stupid.

That's exactly the way UK courts are heading, see Getty vs Stability AI. The court ruled that there's no infringment because the model doesn't store exact copies, just derived weights, and therefore when it generates new images those aren't copies of protected works.

That's stupid, these courts are stupid

It should have nothing to do with storing copies it should have to do with what the models can produce. And it's clear they can produce copyrighted works, they've just been tuned so they don't.

That shouldn't satisfy anyone.

> they can produce copyrighted works, they've just been tuned so they don't.

In other words... they can't. A different one can but this one can't. The court is not stupid, and will consider this fact.

They can't unless the tuner produces a copyrighted work for whatever their purpose would be, because they can step in and "detune" the thing at 3am for a competitive advantage.

This assumes the "tuning" (I assume you mean fine-tuning?) is lossless, which it isn't.

You can produce copyrighted works and have just been tuned not to lol. I fail to see how limiting an ability to comply with the law is any different from just complying with the law?

Proprietary Computer systems shouldn't get the same leeway that humans do. Simple as that.

[deleted]

The point is that you can't steal someone else's content 1:1. But you can use it for a different use (say, display the tweet in an article, then comment on it).

except a bunch of paywalled stuff did end up in training corpora

Pretty sure it's legal when you do it for your own use (same as browsing a website) but it's illegal to redistribute web scraped results.

It's not that cut and dry or else search engines wouldn't be legal. It depends on how much is used, for what context, etc. This very well may wind up being fair use.

Search engines "modify" it ie. show snippets + direct to the actual site.

In general fair use pretty much always requires it to be transformative and/or point to the source. Simply scraping it to prevent people from going to X isn't free use in any definition I've heard.

Remember ~10 years ago when Google had "cached" versions of the websites? I believe they removed that feature for this reason.

Now I want to know, what happens if you redistribute an "AI summary" of the copyright material?

Thats "Billionaire Use" -- its like a "Fair Use" exemption but for billionaires.

Aaron Swartz died so that "AI" billionaires can live.

[flagged]

Someone should add a hitman to the corner of the image

scraping content is mostly legal, redistributing content is not.

if you started doing the same to, say, instagram content both meta and individual creators would sue you as well.

sites like archive.ph are in a similar bucket btw, and yet nobody's complaining (except websites seeing people evading their paywall). but at the end of the day it's not really fair to apply laws differentially on the basis of whose political ideas we like more.

The copyright implications of displaying Instagram content elsewhere are still unclear. There were many lawsuits over Instagram embeds. https://copyrightlately.com/legal-embed-instagram-photos-web...

Meta has no exclusive rights to the content on Instagram, and X has no exclusive rights to the content on X. They have a non-exclusive license to republish it, etc.

>scraping content is mostly legal, redistributing content is not.

So, I take it some Twitter users took an issue with their tweets being redistributed by another platform and sued XCancel?

Because surely you're not claiming that Twitter has any ownership of what gets posted on that platform, are you?

>if you started doing the same to, say, instagram content both meta and individual creators would sue you as well.

By that logic, Meta could sue me for posting my photos on other platforms.

Meta doesn't have a case here, and neither does Twitter, but something tells me Twitter is involved in this nevertheless.

It’s the rehosting of the content not the scraping

Edit: not a moral stance

I love how the content belongs to them when someone else reposts it but it belongs to the user if the content is illegal. Such a double standard with these social media and AI companies. Why do we put up with it?

> Why do we put up with it?

You let it happen. Once people stop letting it happen, it'll stop. But social media is apparently the new "opium of the masses" so here we are and no one wants to do anything.

I actually do want to do something. I started to scrape Twitter to make it freely available. The web should stay open.

I am also doing something, by hosting one of the public instances of Nitter.

Nitter is useful for sporadic random access to tweets, but for public feeds like municipal authorities etc. it would be useful if someone scraped the feed and re-hosted the feed from their own server, without being hobbled by rate limits. Is that what you're doing?

Yes, exactly. I limit the scraping to popular enough accounts. The wikimedia community is already organized to decide who deserves a Wikipedia page so I use that.

Please do share with the other instance operators in #nitter:matrix.org!

We put up with it mostly because it enables a lot of communication to happen at all.

You want to hold the corporations legally liable for the content their host, you can say goodbye to basically reddit as a whole, any twitter clone, youtube comment sections and a whole more stuff.

Did the users of X consent to XCancel copying their posts to their servers?

Everyone who goes to twitter copies the posts to their computers.

Spot on. This is where a lot of these "terms and conditions" break down logically. Viewing some content on the internet is literally copying it.

So is the distinction that xcancel served the content? But when I run

  mtr xcancel.com
I see a bunch of hops between me and them. Every one of those hops is literally copying and retransmitting all the content. Are they not also serving it?

No, this is where programmers rules-lawyer in ways that actual lawyers don't and then get law stuff hilariously wrong. No judge thinks that viewing an HTML page is downloading it, because downloading means saving a copy to your computer, not just looking at it. Even having an internet cache folder doesn't count as downloading. Even copying the file from the internet cache folder to somewhere might not count as downloading, although it'd still be a copy.

Same as when LG said their TVs don't record you and then Hacker News said "how can they detect voice commands if they don't record your voice"... facepalm.

I think its more than the law doesn’t understand the technology.

It's definitely that programmers don't understand law.

I don't pretend to understand law, mostly it just doesn't make sense at all.

It makes more sense when you remember it's not a computer program and the things that are written in the law are not the things that will actually happen in the way that "if(foo) bar;" makes bar happen if foo is true. It's more like a book of excuses you could use for why you didn't do your homework.

Then the other side also has to bring an excuse for why you were supposed to do it, and if the principal thinks their excuse is better than yours, you get detention.

If you tell the principal "I don't have to do my homework because work means employment and it's illegal to employ a minor" you'll get detention for not doing your homework and extra detention for being a smartass.

And this example is not just due to people not taking the trouble to write fully specified rules. I don't think such rules could even be written. You can just do your best to cover the cases you can think of. The complexity of society is incomprehensibly vast and constantly changing, and the law has to have wiggle room to account for it.

You don't want fully-specified rules because a rule with strict boundaries has loopholes. You actually want a clearly allowed area, a clearly disallowed area, and a gradually increasing gradient of punishment in between, so that a small change in behaviour produces only a small change in punishment, and avoiding punishment requires a large change in behaviour.

Exactly!

Could you elaborate in what way you find the law mostly doesn't make sense? It has to be flexible in order to work with actual humans. Why should visiting a page on your computer count as copying? Usually when we talk about copying it's someone making a duplicate so it can be accessed later. Only a very technical user is going to be diving into their cache to view that content after the fact. The vast majority of people don't understand that the browser is storing anything on their computer, much less how to access it before it's purged.

I can't remember the court case, but Blizzard did argue and win in court that WoW Glider's producers violated copyright law. If I recall correctly violating the TOS meant that an unauthorized copy made by executing the file chasing it to load WoW into RAM was created.

It looks like that was MDY Industries, LLC v. Blizzard Entertainment, Inc., which relied on MAI Systems Corp. v. Peak Computer, Inc. for the relevant part of the ruling.

The person I was responding to was saying that anytime you viewed copyrighted content with a browser you’d necessarily be committing copyright infringement. I’m not a lawyer but I can imagine that the reasoning there would be slightly different from someone simply viewing a post in a browser as part of the intended use of the site.

Oh yeah, I understood your point, but given MDY Industries, LLC v .Blizzard who knows what the "right" judge would rule? With IP laws these days we're really getting into weird places.

Yeah I gotchu. I have no idea how that would go in a court for real. I was just trying to explain how and why things are the way they (sometimes) are

> Why should visiting a page on your computer count as copying?

Because there's no physical mechanism for the information to be transmitted over a computer network other than by copying the bytes.

Note this is distinct from broadcast systems like analog television or radio. Packet switching networks only function by copying information and storing multiple copies around the internet, including in your computer's RAM (and disk, if cached).

So a legal definition that says "this kind of copying is copying but that other kind of copying isn't copying" makes no sense at all. Like many other legal definitions--it's all about what has been successfully snuck past a jury at one point or another in the past, without any heed for how things actually work.

It's not about "how things actually work", the law is there to regulate human activity. The law tends to call these copies on the wire, in RAM, in caches, etc. "transient copies", which is fine until a human starts using them as non-transient copies, e.g. saves them for later.

You could argue that your MP3 of Enjoy the Silence is actually just a big number, and you can XOR it with 0xFF and it's a completely different big number, and you just happen to XOR it with 0xFF when you want to listen to it. The courts would look past that, and instead determine if you created that "big number" by MP3-encoding the track from a CD you owned (legal), versus obtaining it from some file-sharing network (not legal)

Classic essay about techies not understanding the law: What Colour Are Your Bits? https://ansuz.sooke.bc.ca/entry/23

> Because there's no physical mechanism for the information to be transmitted over a computer network other than by copying the bytes.

Your response seems to ignore everything in my comment other than the second sentence. I was asking why that detail should matter as far as the law is concerned, and I gave some reasons I don't think that would be good or practical.

There's the matter of linking to copyrighted works: https://en.wikipedia.org/wiki/Copyright_aspects_of_hyperlink...

If your link is set up to make the image display immediately (that is, you wrap it in image tags, or as in one case, embed Instagram posts) then you may be violating copyright. What's more, in Europe, just a hyperlink to a copyrighted work violates copyright.

Conclusion: copyright is not about copying, it's about access.

Sure, but that seems different from what I was addressing. The person I was responding to was saying that the law as a whole usually doesn’t make sense. They were saying that in the context of arguing that if the law didn’t consider viewing a page of copyrighted copying as involving copying due to the technical basis of it having to transfer bits to your computer then the law didn’t make sense. My point was that laws don’t have to encompass or fully specify all edge cases, and that the ways laws are written can be open to interpretation. I think I removed a sentence before posting about the purpose of finders of facts in the US system like juries or judges in bench trials.

> Your response seems to ignore everything in my comment other than the second sentence.

I deliberately ignored it, because it was all irrelevant.

> I was asking why that detail should matter as far as the law is concerned, and I gave some reasons I don't think that would be good or practical.

I have no idea at all how anything should matter as far as the law is concerned. Not my problem, unless I somehow get caught. But not getting caught is a problem grounded in reality, unlike legal ones. I think I can manage that.

That said, if laws about computers don't comport with how computers actually work, I'll take extra amounts of glee in violating them.

And, even more gleefully, nobody will be able to detect my violations. My internet traffic will look identically the same as someone "innocently copying" or whatever.

> I deliberately ignored it, because it was all irrelevant.

Attempting to have a discussion with someone who participates like that pointless.

i just screenshotted this comment without your consent

I think users implicitly consent to their public data being public data when they put it in public together with other public data.

How I view that public data they decided to make public data, is none of their business.

i have not deputized twitter to post takedowns on my behalf. its my content and copyright, not twitter's

Why can't they distill twitter when AI companies distill everything including twitter? Distilling is copying and redistributing.

Are they distilling?

Distilling isn’t copying and redistributing, for the same reason that you reading a story and then writing your own story based on ideas you learned is different from you reading a book, writing all the words down verbatim, and then publishing it as your own.

That's not what distilling is either. Distilling is training your AI to exactly copy someone else's AI.

Because LLM's output is not copyrightable?

No it’s not

How do you know data is copied to their servers?

> If you serve as a mere conduit for automatic transmission of user communications, there are no other qualifications or obligations you need to meet. If you serve a caching function, in addition to the two requirements above, you must maintain comply with the notice-and-takedown process.

https://www.copyright.gov/512/

https://internetcases.com/2024/02/12/dmca-subpoena-to-mere-c...

That is what all LLMs could do in 2023, verbatim, before they trained it out of them in order to keep up the pretense that there is no plagiarism. Now they all obfuscate the original or refuse to cite.

There's a button on X that allows me to repost someone else's content.

there's a setting to disable that

That’s not rehosting

Edit: iPhone autocorrected my OP which meant to say rehosting not reposting

Grok can do that if you give it an HN thread or other websites.