The legal ground is: If you're rich enough you can do it.

Now only big tech companies can train models

AFAIK this does not set a legal precedent as it has been settled and last summer finding is that Anthropic was wrong for "acquiring books illegally" not for training which is fair use.

With model distillation being so effective now nobody actually needs to pirate books to train their models. You can get an open-weight Chinese model and get all that. Or you can just buy the books or buy a library - there are many creative solutions here that aren't piracy and not going to cost you billions of dollars.

The moat right now seems to be the compute resources which might actually be worse for us common folk than a legal moat as we need compute for many more things that aren't LLMs too.