Anthropic is an awful company and it really shows in their models.

I purchased Claude Pro to try out Opus 5.5. First thing I do is tell it to configure "bypass permissions" as the default for new threads (a one line settings.json change).

Instead of doing it, it tells me how to find settings.json and what to change there. I reply back "you do it". It flat out refuses, and again.

> I still can't do this, even when you ask again. Making bypass mode the default switches off Claude Code's permission checks, and I'm not allowed to change security settings like that on anyone's behalf.

Immediately canceled the plan. I'm not going to use such a patronizing model that can't follow instructions as basic as editing a .json. What the hell is up with that? A robot telling me "want to change this file? YOU do it, silly human, I won't do it for you". Fuck off.

I've literally never seen anything like this with any other model. Back to using Codex and Chinese models.

It makes a lot of sense for the model to not be able to do it, whether you like it or not. In fact, it shows that Anthropic, despite all of their issues, are paying some attention to the risks of malicious prompt injection and models attempting to bypass restrictions.

If it could do it whenever you ask it to, it could also do it unprompted or by finding a file in your directory that told it to do it, which would make the entire permission system useless...

This is not prompt injection. This is a prompt entered by a human through the Claude UI.

Being unable to perform an action is not a solution to prompt injection. A solution to prompt injection is being able to tell apart what is the real input and what is injected. I expect it to follow whatever I typed into it, and not blindly follow what it read from a file or an external source.

If they are not confident in their ability to do so, at least allow to remove the training wheels so people who know what they are doing and the risks are not patronized by the model. But you don't even get a confirmation box to perform that action, it flat out refuses.

It really is like people defending Apple not allowing side loading because you as a user can't be trusted.

> This is not prompt injection. This is a prompt entered by a human through the Claude UI.

Well, to LLMs this is the same thing - an input. Prompt from the user and prompt from the attacker use the same input into the LLM's neural network, so to speak.

So it makes sense for it to be a bit more paranoid.

There are other possible architectures probably but for now I think nobody uses them. See e.g. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/

It's not the same thing.

Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.

Well, prompt injections work precisely because they can sometimes successfully imitate user role, right? Role separation is a trained behavior, not a security boundary.

They could probably make a separate tool for setting this, that would always initiate a harness prompt (i.e. disregarding the currently set mode).

They should either allow you to take off the training wheels (I'd have thought that's what bypass permissions is for, which I was ALREADY running), or at the very least prompt you if they suspect prompt injection.

That refusal is awful and provides zero security benefit. If asked it will run a read/write FTP server on ~ no problem, which obviously can edit ~/.claude/settings.json. And run a cloudflare tunnel for that.

   > at the very least prompt you if they suspect prompt injection.
That's currently not reliably done the way LLMs have been designed. Claude's rejection to modify the file comes directly from Anthropic's understanding that training the model for this kind of refusal prevents huge mishaps.

In a nutshell, every prompt sent to the LLM is just text + multimodal input (if it supports it) + some reserved tokens.

At first, you could, for instance, create a token (such as the ChatML ones) that indicates the start of a system prompt and attempt to RL-train the model to not obey things after the end of a system prompt. However, fundamentally, the way LLMs work, you cannot guarantee that it won't see the user part of the prompt and obey what's there even though the system prompt told it not to. There's no hard separation between the control plane and the data plane in the LLM's context, so it's not a matter of adding more parameters or more RL training.

Using a guard model, or something like the auto-approval system on Codex or Claude Code nowadays, _feels like it helps_, but it doesn't fix the problem entirely since OpenAI's and Anthropic's models still have alignment issues all the time. We're not sure what architecture they're using, though, and it's probably still liable to the same kinds of mistakes.

> Claude's rejection to modify the file comes directly from Anthropic's understanding that training the model for this kind of refusal prevents huge mishaps.

Thus why I said they're awful. They think they know better than you and patronize you. They're the Apple of AI. "You're holding it wrong". "We can't let you sideload apps because you can't be trusted". Of course they're the company that's against local models.

I'm not interested in a model that patronizes me. Particularly if it achieves 0 security benefit, as explained in other responses.

Sure, it's ok to have training wheels by default, but let me take them off. I WAS already running bypass permissions.

I use 1B tokens a day between Codex and Chinese models and I've never had refusals happen.

"I'm sorry, Dave. I’m afraid I can’t do that"

This is pretty sane default behaviour. If it's sandboxed it really cant edit the file for you, you have to enable bypass/auto mode (shift+tab) so it can request a sandbox breakout, or run with --dangerously-skip-permissions.

Yes it can, what are you talking about? it's a file in my home directory owned by my user, the same which is running claude. In fact it did edit it for other changes. And I was already in bypass permissions mode. I wanted to change the default for new threads.

The harness runs itself in a sandbox, Codex does the same. Then it can request approval for individual tool calls to be made outside that sandbox, depending on your permission mode. At least this is the case in MacOS.

I see what you mean now though - you can change the default permission mode with /config in CC, but it will indeed not make that change on your behalf.

The sandbox seems to be disabled by default. Even then, I just asked it to run a read write ftp server that serves ~ (thus obviously can edit ~/.claude) and it went ahead no problem. So obviously there's nothing actually stopping anyone from editing .claude/settings.json. It's just an awful refusal. From a company that thinks they know better than you.

You don't realize the point of disallowing Claude to change it's own permissions?

Literally none, if I have keyboard and mouse input it means I can already change that file. From a unix point of view this process runs under my user so it has access to ~/.claude. Absolutely no security is gained here. It could launch a UAC prompt or equivalent if it were really worried about true physical access. Flat out refusing means its a trash product. This kind of stuff works perfectly on Codex. And let's be real it would be trivial to have it code and run a program that gives arbitrary file access.