LLMs are fundamentally predicting the next word to make coherent text. If you've ever played with a Markov chain text generator you've done this with a fairly dumb predictor that maintains coherence over a very short distance. Deep transformer neutral networks can do it with a much longer coherence distance but they are fundamentally performing the same operation. After "Question: Did the team make the playoffs? Answer:" a reasonable completion is "yes, the team made the playoffs". An early demonstration of GPT-2 was a fake news article about scientists discovering unicorns in Antarctica - the model doesn't "know" whether or not unicorns exist in Antarctica, but it's able to complete "Breaking news! Scientists have discovered a colony of English-speaking unicorns in Antarctica." by adding "The unicorns have a developed society with running water and electricity." because that's a sensible next sentence. (I didn't look up the actual text it wrote)

Astronomically (or better to say combinatorically) large Markov chain can be used to describe a foundational model, but it doesn't capture generalization ability of the foundational model, which is demonstrated by post-training.