> both outputs of Claude Mythos, their (still) unreleased advanced model
That sentence gives the impression that Mythos might be released in the future. That's clearly not going to happen - it's already "released" in as much as selected, trusted partners can access it, and the rest of us get it in the form of Fable - which is Mythos but with filters that downgrade you if you try to use it for anything even remotely related to cybersecurity or biology.
(The other day Fable 5 downgraded me to Opus after I asked it to explain the difference between tusks and teeth.)
Genuine question here, why would you ask Fable to explain the difference between tusks and teeth? That's a task that can probably be handled by Haiku.
It's the default when I pop open the Claude iPhone app, I usually don't bother to switch it.
because as OP of the article points out, it's not a guaranteed hit with these models no matter how capable. they are stochastic (but not parrots), humans are too (but in a different way), and so there is a chance that the answer is wrong. If I wanted the most accurate answer I'd be confident in believing, I'd ask the most advanced model.
Now obviously you can and should retort with hallucination and confabulation rates from external and Ant's own reports per model (pretty sure more advanced models are good at lying better, not less) instead of going with dumb "more expensive more accurate" mental model, but general principle stands for me still AFAIK.
it's not exactly rational I admit, but if I'm going to base my own work and reasoning from an LLM I'm going with the best available. This seems to trip up most normies because they are too lazy or too greedy to pay up for premium access and see for themselves why most of us are both awed and afraid. Generally, I'm too biased and too deep in ML/DL cargo cult (been in it since 2016) to know if the skepticism and disdain for such usage is warranted.
In general, I think the tools are broadly toxic in a Dune-sense of making me think less for myself, because just as any HN-poster knows coding and doing mundane low-level stuff is necessary the same way doing stretches is necessary before any workout. The process itself is what keeps your brain strong and its gradients from veering into overfitting. I'm not overly bullish on the whole reaching for the stars ending with these things. Paradoxically, you using them eventually hobbles both you and the model, because you become dumber and then you bottleneck their ability to self-direct (broadly true for next Mythos/GPT-7).
Sorry for a long rant, was just anticipating some things I'd have to say for myself.