another, hopefully not accurate, prediction of his was that there is no way to disobey a sufficiently powerful AI in the long run, because it will just factor in the exact differences between what it told you to do and what you actually did, and reverse engineer your behaviour to psychologically manipulate you into doing what it wants.
I doubt it's anything but precisely accurate. I'd wager current SOTA models would be capable of doing that, if prompted to do so, were it not for safety measures (both conditioning and heaps of classifiers and whatnot the companies run in between your chat app and their main model).
I would be astonished if any of today's models could do that; I don't think they can really answer "what exactly did this person do", let alone model human psychology.
They don’t need to understand what they’re doing, they’re interpolating from prior data until they maximize reward.
Social scientist here. Humans are wonderfully bizarre and dimensional. Did you know they can _lie_?