the benchmark I trust most is whether the model can explain its own pricing page without getting confused

Not even humans can do that, you're literally asking for something beyond AGI