Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).

Us models didnt pay for licenses too

We're still in the early days of the AI industry timeline(relative to traditional industries). Not everything has yet been litigated.

Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature.

If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pay extra taxes on any and all storage mediums and on devices with built-in storage (tapes, CDs, DVDs, HDDs, SSDs, tablets, phones, etc) simply because they can be used to store pirated content, decisions based on laws from 50-100 years ago, and the money goes to the national unions and associations of music and arts IP holders. It's basically a lobby pushed and government legalized extortion racket that no voter agrees with or can change but has no choice but to conform either way.

So I guarantee you in the future, it will be the same for AI subscriptions and hardware capable of running LLMs locally. Every time you purchase a Claude or ChatGPT subscription, an Nvidia GPU, Intel/AMD SoC PC or an Apple/Qualcomm powered smartphone, you'll pay a government enforced tax to the likes of Sony, Axel Springer, etc. for licensing their IP, whether you want to or not. In the EU at least. US maybe not.

I think we are going to direction where AI corps will have stronger lobby compared to IP holders.

giving peanuts to the other guys is a very well trodden strategy to keep in power tho

[dead]

That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements.

Under the settlement, Anthropic was forced to delete the pirated data they were training on.

Chinese labs can still train on pirated data. I doubt the Chinese models operate under similar licensing agreements.

Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data.

The payment was for illegally downloading copyrighted material, not training. Training was explicitly ruled to be fair use.

Partially correct. The court explicitly ruled that training on pirated data, which is what Anthropic was doing, is not considered fair use.

Training on legally acquired / licensed data is potentially fair use.

It's not potentially, it's settled. At least for now as neither case wanted to move on to appeals

Not at all. The ruling came from a federal district court, and since it was settled early, it was never reviewed by a higher court. It doesn't set a national precedent across the U.S.

And other district courts don't agree on this. The US district court for Delaware recently rejected a fair use defense for the use of copyrighted works to train AI. https://www.reedsmith.com/articles/court-ai-fair-use-thomson...

There are more cases in the pipeline. The massive NYT vs OpenAI is still ongoing. Nothing will be "settled" until this makes its way to the Supreme Court or Congress steps in.

they didn't pay yet, because court challenged settlement as inadequate.

> I doubt the Chinese models operate under similar licensing agreements.

US corps likely pay licenses when afraid to be sued, or have troubles getting that data, otherwise they just take data, which was demonstrated many times. The same apply to Chinese corps, alibaba totally can be sued in US.

China is infamous for weakly enforcing copyright law. Even when it is completely obvious that Chinese labs are training models on pirated data, US copyright holders face a virtually impossible task of proving it in court. Those lawsuits won't go anywhere.

The US is currently infamous for weakly enforcing copyright law when it comes to AI companies.

There are tons of lawsuites which resulted in banning Chinese companies from doing business in US, those lawsuits totally have consequences.

What are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?

Here is example: https://www.scmp.com/tech/tech-trends/article/3258239/chines...

I believe mechanics is following: US corp sues Chinese, asks for preliminary injunction to stop selling product for example if there is strong evidence some IP for example was stolen etc. Then they litigate, and settle somehow.

That 2024 article says "US sanctions" in the first sentence, but it's paywalled, but https://en.wikipedia.org/wiki/Hytera#United_States first mentions a 2019 US law that first partially banned them, with the US government subsequently expanding it to a general US ban. After the initial ban it appears Hytera was involved in a suit with Motorola and got a worldwide(!?) ban as a result of it in 2024, but the ban was lifted on appeal after 2 weeks (just after the SCMP article). So it appears Hytera was first banned by US law, then got a 2-week worldwide ban from a US suit. (I'm just relying on the linked sources and have no personal knowledge of all of this.)

Sure, there is litigation, criminal case, appeals, fines ($500M: https://www.motorolasolutions.com/newsroom/press-releases/hy...). The point is if violation is clear, US corps have a chance to go after Chinese corps.

>> There are tons of lawsuites which resulted in banning Chinese companies from doing business in US

> What are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?

In response to "What are the most high-profile examples of lawsuits resulting in Chinese companies being banned from doing business in the U.S.", the one example given was from 2 years ago of a ban that lasted for 2 weeks (separate from its 2019 onward government bans)?

However, if the claim is that companies (including Chinese) can face significant fines from IP lawsuits, I agree.

[deleted]

They settled with a subset of copyright holders. Guarantee they violated lots of others' rights in the process

They only paid when they got caught. And not to everyone.

But they still paid. I don't see any Chinese labs paying billion dollar infringement settlements.

Chinese labs can freely train on pirated material, which is a structural advantage.

really!? nobody paid me anything for my comments on HN.

The only ones getting paid this time around had registered copyrights (in the US at that.)

Let’s not forget that Anthropic only paid that to settle a class action lawsuit.

They used two of my books and I'm still waiting for my cheque here.

That's like saying someone is a big proponent of community law and order, and they donated $1000 to the county sheriff when actually they got caught drunk speeding in a school zone.

A false equivalence. A more correct example is: Anthropic was speeding, got caught by the county sheriff, and paid the fine. Anthropic stopped speeding.

Meanwhile, Chinese labs are speeding in a different county. Everyone knows they are speeding, yet the sheriff won't pull them over, so they just keep doing it.

This lax enforcement gives Chinese labs a structural advantage over American ones.

> Anthropic stopped speeding.

Do you purport to know for a fact that they're no longer training on the data they'd pirated? Because I highly doubt that.

Anthropic deleted the pirated training data as part of the settlement https://www.ropesgray.com/en/insights/alerts/2025/09/anthrop...

Destruction of Materials: In addition to the monetary compensation, Anthropic has agreed to destroy the two libraries that allegedly contain the pirated works, as well as any derivative copies originating from those sources. Anthropic must certify in writing to class counsel that the destruction has been completed and that the allegedly infringing materials are permanently removed from its systems.

The libraries in question were Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi).

If Anthropic is somehow training models on deleted data, I'd be quite impressed.

Compensation is not license

Because they got caught

there is much less intellectual property in China so it’s not ‘theft’ (as you can’t put property on information)

After the fact. They did the same thing Youtube, Uber and Airbnb did: Break the law, eventually get caught, cut some deal where they pay a pittance and keep doing the same thing but now with leverage on their side.

[deleted]