Sounds reasonable... That the model is sometimes learning a lossy vector representation of something symbolic in nature... Sure, a NN can approximate a function?

They say this holds in... Some examples they found?

I don't enough about this area