If we could come up with a system to classify the probabilities across a large number of candidate words (or components thereof) then this could actually be good at producing text, one element at a time. We could call these elements 'tokens' and picking the right one could be called something like 'decoding'. Crazy idea but hear me out...
On a more serious note, it will be fascinating to see how this different spin on modelling inference will create new paradigms or slot into existing ones.