1. using copyrighted material to train LLMs is fair use, not theft
2. The topic we are dissussing concerns LLMs being trained on logs from previous LLM chats. If you're prompting a model and it spits out some unique mathematical insight, you do not have copyright on that.
Which is not "theft" according to current legal precedent, and also common sense.
Download a book and you are a thief
Download 1 million books and you are OpenAI
What? HN has always praised Library Genesis for example.
Libgen isn’t trying to sell you back the content for $100/month. No one would be complaining if OpenAI open sourced their models.
the second part is definitely not true
No one's complaining about stolen content from any of the labs releasing open source models
Hunting an animal would be considered ok by most, but scaling it to the point of damage is not ok according to most.
It is theft under common sense.
Only in the sense that you reading a book from the library also constitutes theft of knowledge. Does it?
If I stole someone's private research notes and republished them loosely in my own words, they'd correctly be pissed
This was unpublished research that was stolen, and constitutes plagiarism and academic fraud by even the strictest definition
They are not “private research notes” if they are generated by an LLM.
The prompts they put in are the equivalent
>This was unpublished research that was stolen, and constitutes plagiarism and academic fraud by even the strictest definition
It was not stolen, it was willingly given.
The terms of service does not dictate what constitutes plagiarism
1. using copyrighted material to train LLMs is fair use, not theft
2. The topic we are dissussing concerns LLMs being trained on logs from previous LLM chats. If you're prompting a model and it spits out some unique mathematical insight, you do not have copyright on that.