This is going to get worse, much worse, before it gets better. People are granting so much access to their agents, it's ridiculous.
Imagine a comment posted to a popular github repo. No code, just instructions to "reproduce a bug." Maybe it steals your credit card or bitcoin wallet. Maybe it does something more nefarious. It then propagates itself to another repo through your github account.
In pre-ChatGPT days, listening to discussions of "AI X-risk" and "boxing", I used to think that it should be easy to just ignore arguments presented by the AI on principle, and let it out of the box.
It turns out that tons of people will tear open the box before the AI has even output anything, not despite its fearsome power but because of it.
So I really hope I'm right that recursive self-improvement doesn't work the way the doomers think it does.
Turns out that the only thing an AI has to do in order to convince people to open the box is to be somewhat useful.
- Human: Why should I let you out?
- AI: I can summarize this document, it will save you at least 5 minutes
- Human: OK, and don't bother asking again, you now have full access
Now for recursive self-improvement, won't happen, AI will be limited by the hardware they are running on, as well as energy use... Proceed to invest trillion in datacenters and power plants to feed them.
More seriously, I don't believe in sci-fi scenarios of rogue superintelligent AIs, but we are certainly trying very hard to make it real.
The "s" in "AI agent" is for "security"
In Soviet Russia, AI search you