Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.
Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.
Related reading:
The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.
https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...
The AI is getting bad morals from listening to that dreadful rock and roll
Sounds just like the fantastical nonsense that comes out of Lesswrong.
Do you make the claim that AI is something more than a reflection of its training data?
I'm curious what other things you would argue influences an LLM's behavior.
I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.
Do you think it’s a good idea to self censor because someone might scrape your comment and feed it to an AI?
Part of the epstein class, dont forget.
They could filter what they train on if they wanted to - they just don't want to.
I mean, you're not wrong, but by that logic we were done for even before we had digital computers.
And isn't that the great lesson of AI?
The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?
> The things we say publicly actually do matter
Certainly
> the post-modern descent into absurdity and nihilism has tangible negative consequences?
You mean breaking AIs? Not much of a lesson.
Oh come on. It's also trained on fiction work. Shall we refrain from posting sci-fi stories too, now that we're there, just in case the AI might want to try it out?