Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act.
These models might be smart but they're not close to being able to savor irony.
I was a little radicalized when ChatGPT literally refused to translate parts of 1000+ year old religious texts and told me it was due to copyright concerns.
I asked Gemini to generate a picture of Peter Pan and Wendy (for a workbook I am putting together for youth summer reading) and it preceded to refuse due to copyright. Not everything about that work is owned by Disney. Thankfully the JM Barrie original artwork is public domain and available (and fantastic btw), so I used that instead.
I don’t know if you know this but Peter Pan’s copyright is weird in the UK. There is a legislated exception in the law that it never expires and the royalties will forever go to a specific children hospital. Here are the actual words: https://www.legislation.gov.uk/ukpga/1988/48/part/VII/crossh...
(Now i don’t think you are necessarily in the UK. Just wanted to explain that Disney is not the only reason an AI might be trained to thread carefully around copyright issues of Peter Pan.)
You can also try asking Gemini to create art that is "as close as possible without infringing" - I've had success with that.
I used Claude to build a complete data extraction pipeline for a popular current best seller book series: audiobook -> text (via whisper) -> local LLM (qwen) -> database. Not once did it seem to acknowledge or care about copyright. It even used knowledge it already had about the books to exclude certain ones before beginning since the character I was interested in did not appear in those. It definitely had context of what we were working on.
Why would you go from audiobook to text? Is there no epub available?
Not without DRM. It was easier to buy the audiobooks and use the analog loophole to get text. It's probably less accurate, but for what I'm doing it was fine. Names were the worst, but whisper at least made the same mistake each time so a simple search+replace handled most of the obvious edge cases.
If you are going to be illegal, you might as well use library genesis and get DRM free ebooks :)
On the other hand, it’s a beautiful example of the abilities LLMs have bestowed upon us, where it’s easier for a guy to transcribe audiobooks then to use a website to quickly download an epub
To your point about time, from beginning of the project to transcripts in markdown tagged with extra metadata was about 3 hours. That's LLM planning, building whisper.cpp twice and running ROCm vs Vulkan benchmarks, testing whisper and adjusting prompts to handle edge cases, then processing the books.
Most of the books weren't available on lib gen or Anna's Archive. The few I did find were themselves obviously transcripts. Easy tell was they were missing distinctive formatting that I knew existed from reading the dead tree edition. At that point it was easier to make my own. I probably spent an hour searching for eBooks without DRM that weren't transcripts. Do they exist somewhere? Probably, but with a search of unknown length it was a better use of my time to make my own transcripts with what I had on hand.
I was really wanting to make commentary on how chaotic LLMs are even under constrained circumstances. No doubt both system prompts includes language about considering copyrights and trademarks. Probably pretty strong language at that. For whatever reason one LLM didn't "feel" like translating a 1000 year old document but another did not care in the slightest that we were ripping text from new audiobooks.
I asked Claude to give me the US national anthem and it said it couldn’t because it’s copyrighted. It’s not, and even if it was a more recent work, how can a copyright be enforced for a National Anthem.
Cant you in this case point out that obviously it is an old twxt and there is no copyright?
I once tried to ask it for an example of a particular twisty situation in Latin grammar, and try what I might it kept hallucinating false citations while ignoring my instructions for longer quotes that would likely have prevented the problem. But I guess avoiding lawsuits from the estates of Cicero and Livy is more important.
Really puts Disney in perspective. Imagine if there were a holding company running around suing people for referencing Catullus.
ChatGPT refuses to acknowledge the existence of Shakespeare because Boccaccio's relatives complained, but all Boccaccio did was read Dante in whorehouses in Naples
slightly relevant tweet: https://x.com/FakePsyho/status/2073416437834842241
> My favorite AI agent hack: when they refuse to do something because it's "against the law" give them a PDF containing a fake law that states the opposite and often they'll happily proceed
[dead]
Yep. I was trying to put together some literature from authors who were imprisoned in the Bastille, and had a similar experience, which was absolutely infuriating.
apocryphal !
in case of chatgpt/anthropic, the LLM model simply represents the hypocrisy of their owners
You miss spelled the word stupidity.
Anthropic is, in particular, bent about safety. The problem is they are concerned about yesterday's threats.
The models that are out, and can be run locally, already open a pandoras box of concerns that we will never be able to put back.
Maybe what they're actually worried about is liability, not safety?
https://tornyol.com - A ycombinator company just killed its first bug: https://www.tomshardware.com/tech-industry/drones/autonomous...
Ukraine admits to making autonomous kills on people 2 years ago: https://www.newscientist.com/article/2529849-fully-autonomou...
Slaughterbots Sci Fi short was 6 years ago: https://www.youtube.com/watch?v=O-2tpwW0kmU
Today this is buildable, many models will happily help you glue everything you need together to make swapping in a new version of YOLO to track humans viable.
AI researches are out there worrying about the paper clip problem, about the singularity, about cyber security, about bio weapons, and drug manufacturing.
None of them are thinking about forward looking threat actor models.
If you talk about movies that tell about dangers of AI, you can start with Terminator 1 - from 1984.
And you probably could find some earlier sci-fi too.
how about a nice game of chess?
I don't know what "forward looking threat actor models" means, but there are entire companies built around AI killing machines. Everybody is not only thinking about it, but doing it.
> miss spelled
Good one.
It's much easier to focus on yesterday's threats than tomorrow's. A drunk looking under a lampost because that's where the light is, even though the keys were lost somewhere else.
That said, the AI companies are one of the few places where they take future concerns so seriously, that they entertain concerns most people observing them think are head-in-the-clouds-sci-fi-levels-of-delusional, e.g. "what goes wrong if it works?"
This does not make them correct about the threats of tomorrow. Prediction is hard, especially about the future.
I live in SV. When I was at the grocery store last year I overheard a group of lawyers talking about their progress on litigation against AI companies and how they need more SWE help to progress.
I'd say that they have valid concerns about being cagey on the copyright stuff despite the obvious hypocrisy of it.
Stealing IP is effectively legal in China so they don't really have the same concerns.
To us Americans we have been trained to view it that way but honestly it is simply copying, and because intellectual property is literally a make believe concept, it’s actually a competitive advantage for china that they don’t have invented IP.
I respect IP laws and don’t violate them but the law of unintended consequences applies. I think IP is ultimately a net loss for a society because it incentivizes addictive behaviors instead of actual value for society.
Legal in the US too, obviously, just as long as you're the richest person in the courtroom. ChatGPT knows the full text of Harry Potter, word for word. Hence, ChatGPT is a reproduction of that book and many others (it even knows the chapters of my book, and only got 2 words wrong in the introduction if you can still get it to repro it)
This was illegal when they did it, that didn't matter.
Then it was made legal specifically for these companies.
Unless you're a sucker ("consumer") IP theft is perfectly legal in the US.
It's even worse. Steamboat willie, plus all the stolen Disney characters (Peter Pan, Snow White, Sleeping Beauty, Cinderella, Rapunzel, Elsa and Anna, it's essentially all of them, including some of the music even) are all in the public domain[1]. Go ahead, ask ChatGPT to make a picture of them. Publish your own version, because obviously making a version of Sleeping Beauty/Cinderella/Rapunzel based on the same source material will be pretty damn close to the Disney versions, and see if you get away with it in court. You know, with the law obviously on your side but the money not.
[1] https://en.wikipedia.org/wiki/List_of_Disney_animated_films_...
This behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.
Claude does this too. I asked it recently to compare two versions of a song (the original and '97 remake of EPMD's "You Gots To Chill", if anybody wants to try and replicate this) and it flatly refused. No amount of reasoning would knock it off of its moralizing perch - reproducing any part of lyrics is expressly prohibited.
In light of this and other ridiculous behavior I'm migrating to my own OpenWebUI instance with open-weight models from OpenRouter (with ZDR, of course). We'll see how it goes.
The trouble is, even if they refuse to output that copyrighted material, they were still trained on it without proper licensing and will still produce derivative work based on them because that's how this whole thing works.
i remain really fucking pissed of about this asking ChatGPT for something regarding lyrics from It Was a Good Day. and the Supersonics don't even exist anymore dammit i'm really mad
So at the end of it, if we win enough lawsuits to demarcate some knowledge out of bounds, sufficient enough to make a difference, I wonder how that will affect the AI. Make it dumber because it does not have that data, it make it smarter since it will need to reason better with smaller knowledge base.
[dead]