> LLMs, even frontier models by major providers, still have no reliable internalised way to assess the accuracy of their output
This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.
If your believes were true, this actual quote, pulled from a recent sonnet conversation, would be impossible:
“The XT60’s 15A limit is unsuitable for an appliance that… no, it’s the XT30 that is rated for 15A, XT60 is rated for 30A continuous. For a 20A appliance, XT60 will be fine.”