You could probably keep the Claude slop hidden and have a fresh model generate a paraphrased response for anything human visible and keep the Claude responses as thinking.