I don't understand, if they are only using a subset of the tokens then it's a sparse model. What do you mean by dense?
Could it be some sort of permanently routed MoE where they detect and switch for the whole prompt instead of token by token?
Could it be some sort of permanently routed MoE where they detect and switch for the whole prompt instead of token by token?