It's pretty clear reading from these comments that most HN members have a 2023-era impression of LLMs.
Modern chain-of-thought models with RL post training on verifiable tasks + realistic environments + rubrics are worlds apart from models trained on a simple next token prediction objective.
More money goes into the rubrics and RL environments than individual training runs themselves.
(Yes, at inference-time LLMs still output words one at a time, much like human speakers. But don't confuse the mechanism with the training objective.)
Even with heavy RL post training and rubrics, the model is still fundamentally bound by the next token prediction mechanism at inference. Rlhf and cot just affect the probability distribution of which tokens get predicted next. Take away the heavy agentic scaffolding and external feedback loops, and a single hallucinated token can still derail the entire chain of thought.
But but but....I was told it was a stochastic parrot! I liked that idea because it appealed to my vanity, and it described the gibberish produced by older models with bad prompting, and that was enough for me thank you.
/s