Attention is at about the same level of abstraction as spike trains or action potential, IMO. It’s a mechanism, it doesn’t tell you much at all about how actual concepts get represented. (It merely defines the substrate with which they can be represented.)

The entire field of Mechanistic Interpretability exists because just understanding Attention does not in any way help you to understand why a certain NN responds with a certain hallucination about a certain Chinese boat in this specific context.

> The only remaining thing to study is the data and weights

To me this is like saying “the only thing left to study in the brain is the connectome; probably going to be boring, we understand it already”. It’s almost all of the hard/meaningful stuff! It’s where intelligence and consciousness lives!

The entire field of mechanistic interpretability has gotten nowhere. The formalized, "causal" understanding of latent space conceptualization isn't any better than an LLM cargo cult.

It's wholly possible that you could study one set of weights for decades, and find nothing. There's no guarantee that any patterns outside of human language exist in that data. In this specific context, it's satisfying enough to state that [Chinese] and [boat] were both tokens in the tokenizer, activated by a feedforward pass through weights that favor [boat] after [Chinese]. There's not any guaranteed solution to this. There's not even any guaranteed problem; that hallucination is an expected behavior.