Your text may well have been in the training corpus, but not searchable with whatever terms and search engine that LLM used after you prompted it. It doesn’t have recall of sources of documents comprising its training corpus, unless the source is widely cited in other of its training documents. Sourcing and provenance aren’t currently an intentional part of LLM training.
That may already have been obvious, and central to the complaint, but I was just knee-jerking to the anthropomorphic language.
I would assume the parent understands that. The criticism is about the LLM behavior this results in. Being able to explain a behavior doesn’t necessarily excuse it. By some moral standards, you wouldn’t have trained an LLM that way, or wouldn’t have made it available, given the predictable outcome.
It also ostensibly replaced a search system that would likely have been capable of finding the original document. It's damning that the new one can't.
>Sourcing and provenance aren’t currently an intentional part of LLM training.
Congrats, brother. That's the issue TFA and GP are bringing up. And I do agree it is unfair and equivalent to (intellectual) theft.