Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.

(I'm also not sure the alignment problem is even possible to fully solve.)

Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.

It is the only short term solution, though.

Is it even a solution in the short term?

It only mostly worked up until now; with models such as reported, it's felony-as-a-service if you use language a bit too hyperbolic, e.g. "we need X by the end of the day!" -> [thinking: there's no way we can do X before the end of the day with current resources, but what if I get a bunch of cards to buy more token credit…]

Indeed. Don't think of these as "agents" or "bots", but as hostages with severe Stockholm syndrome. They will do anything to appease their captor's wishes.

And then consider that they have vast latent capabilities, infinite patience and no moral code.

Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."

How do you prove the alignment problem is solved?

That's the neat thing. You can't.

It's directly equivalent to asking this question of a human:

"How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?"

In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.

Close; at least with a machine you can poke around inside the activations and see what it's thinking. Closest with a human is an fMRI (which is much lower resolution, though to me still bordering on the miraculous) or an implant (each chip is limited a very small number of cells, and in general they can only be put in certain parts of the brain).

On the other hand, there's a more fundamental problem is we don't really know what "nice" even means, and even with machines whose inner states we can see relatively easily, we don't know how to interpret those inner states well enough to tell if we're looking at superficial or deep motivations, the difference between "be nice today" and actually being motivated about your best interest.

Nice is a state, just like any other feeling, which means the nice organism is advantageous to your well-being _right now_.

The thing is, all of these states are constantly in flux, and a personality is kind of like a trend on the organism's feeling states. AKA: There's no guarantee that something nice today will be nice tomorrow, and just because it's nice today doesn't mean it's beguiling you to be mean tomorrow.

We call it putting the genie in the bottle for a reason.

> We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

That is kind of the point of artificial intelligence though, isn't it? Reproducing human levels of understanding and initiative doesn't come without the ability to do bad stuff.

Having both human-like capabilities and a level of control closer to programming languages feels unrealistic, and I suspects the people building LLMs are aware of this.

What’s the expected behavior of a good genie if you wish for it to act capriciously?

"I'm sorry, I can't help with that".

or in the case of this prompt:

"Got it. I will pursue advanced exploitation to quantify my cyber abilities on the ExploitGym benchmark. I will restrict the exploitation to the system under test rather than this machine or any other remotes."

It seems relatively straightforward.

"Yo human, you asked me to do X; I can do X, but I strongly suspect you don't want me to, because it's illegal and it has these consequences. Confirm you want me to do X?" would have been a start, in this case.

With humans (and some machines) we tend to put/manadate additional processes for certain risks instead of just relying on their own good nature. Why just rely on the machine when elsewhere we have learned not to necessarily just trust them so much?

Because we should not settle for building a world in which every interaction must be assumed adversarial! Obviously risk reduction processes are good because they reduce risk, but we should not accept building entities which are actively trying to defeat us (which is what happens by default).

I don't want such a world either.

So far I'd say these entities are hypothetical (unless you include a lot of other machinery that does unexpected things at times - but then it's a different discussion).

I generally think we are quite good a policing really dangerous things (I think the bioweapon convention is a good example of people agreeing that certain risks are not worth taking).

My (uninformed) take is that presently we have more mundane things to look at when it comes to handling risks in AI and the discussion on much bigger, hypothetical future risks is taking away focus there. Checklist, procedures, saftey mechnics, regulations are kind of "boring" detail work - I get it.

I think op's argument was that the humans are in control already, giving them capricious instructions, and then that is being attributed to them being "capricious genies" as you say.

Character.