From what I've noticed at least for myself, there are like at least two distinct parts that don't always cooperate. There's a draft part that's random thoughts that just appear and can be converted immediately to speech, but when writing it's actually another part that takes dictation from that one, albeit silently.
If I'm just typing blind I'll somehow end up writing random homophones down with completely correct spelling, like hear instead of here, it's bizarre. So that part is sort of structured autogenerated noise that is then consciously either appended or inserted into random spots in text. Maybe it's more of an LLM first pass then diffusion refinement.