> The monitorability of frontier models is degrading.
Is there more information about why this is happening? Is political pretext because it's what the labs actually secretly want, or is there a real underlying reason this is unavoidable?
> The monitorability of frontier models is degrading.
Is there more information about why this is happening? Is political pretext because it's what the labs actually secretly want, or is there a real underlying reason this is unavoidable?
Chain of thought tokens are vectors that have the same dimension as the input/output embeddings. This allows them to be un-embedded back into text, making interpretability easier.
There is no mathematical reason that the chain of thought couldn't happen in a different dimension. Indeed there are likely many reasons to do so. At this point you'd have to do some kind of (potentially lossy) projection back into the embedding dimension in order to understand what's happening.