> can’t use claude to research AI

What's this about? Where's this rule?

In the system cards. Anthropic will literally make Claude sabotage you silently instead of downgrading you to Opus if you try to use Fable for AI research.

Didn't they walk that one back eventually?

Who knows? It's trivial to "walk back" claims that we can't even verify are happening in the first place. They cannot be trusted.

Don't they just fallback to opus non-silently?

Only if you're doing cybersecurity or biology stuff. If you're a possible competitor, they will invisibly fuck with Fable's weights instead to deliberately sabotage you.

I find it implausible that they simultaneously are so worried about these recursively self-improving systems that they don't properly understand, but can freely "invisibly fuck with the weights" on that level.

Steering Vectors