In the Jev use case, LLMs are horribly uncalibrated. In general, they will not produce good probability estimates.

Their generality also comes with a latency/computation costs.

For the Jev use case for LLMs, do you mean having the LLM produce a probability as text?